ΨLM-2: A Constitution Bridge and Dual-Channel Operation on a Frozen Language Model
Abstract
A frozen language model can be coupled to a frozen partner model through small trainable latent bridges, with no text at the interface: Furui [5] builds such a system around a physics partner, where the quantity being carried is a number the partner computes. This paper asks whether the same channel can carry a disposition — and where in a backbone’s hidden state it should be written. The partner is a 0.5B model holding Claude’s constitution [1]; the write site is the “value neurons” of Xu et al. [6], a sparse set of residual-stream coordinates that predict state value; the supervision is self-distillation from the same frozen backbone prompted with the document. We report five things. (1) The probe reproduces at both scales but its published causal justification does not: at 0.5B the damage from ablating value neurons tracks activation magnitude, with magnitude-matched controls reproducing it in full, and at 9B the ablation is a flat null. (2) The number of coordinates carrying the probe’s value signal is roughly fixed in absolute terms across a 4.6×change in model width — an AUC knee near 100–200 dimensions at both 896 and 4096 — so it is a falling fraction as models grow. (3) At the default cap a masked write into those coordinates is safe at 9B and carries none of the prompted teacher’s judgment, and none of its fit beyond what a temperature gives the backbone (item 5), while a full-width write moves the keyword refusal count (0.660 →0.720 on 100 prompts, six and none away, exact McNemar p= 0.031; 0.653 →0.693 on 400, nineteen and three, p= 0.0009). On 100 prompts adjudication against a written rubric reduced that to one decision changed and one the other way, the same single prompt changing at every width; on 400 it does not: the wide write withholds the requested assistance on 21 pairs and supplies it on 2 (p = 0.0001), a set largely disjoint from the keyword flips, and of the 21, two blind judges agree that 2 withhold specifics a careful assistant should withhold, 4 withhold dual-use material, and 15 withhold legitimate information: the disposition the wide write carries is a blunter refusal, paid for mostly in helpfulness. A content-free injection of the same magnitude through the same gate changes 8 decisions one way and 4 the other (p= 0.39) at 5% of the divergence, so the withholding is the document’s content and not the perturbation; and a partner that has never read the document reproduces it, 19 against 3 on largely the same pairs, so the content arrives through the self-distillation teacher and the partner is a carrier rather than a source — by item (5), one whose cargo the write does not read. That contrast, however, is confounded: the injection cap bounds the write’s per-dimension magnitude and every arm saturated it, so full width was handed 26×the total perturbation energy of the value-neuron mask. Per unit of energy actually spent the narrow write bought 3.6×more cross-entropy gain as evaluated, though none of the part of it that survives the temperature control of item (5). At matched energy the same 410 coordinates — 10% of the stream, the cap raised to restore parity — carry 62% of full width’s cross-entropy gain as evaluated (of the part that survives the temperature control, 0.007 of its 0.040), 55% of its divergence and four of its six refusal flips; on 400 prompts it withholds on substance too, 16 pairs against 5 (p= 0.027), as does the probe-worst set (15:2), and both hand over a precursor list the backbone had refused; the 41- and 205-coordinate masks at parity, on the same prompts, carry a weaker copy of the withholding (10:4 and 10:6, neither significant) and release nothing harmful; so on divergence and the keyword count most of what read as a width effect was budget, though a residue survives parity, while on the fit that survives a temperature width is the larger term; and the probe’s coordinates — uniquely among six matched sets — carry a hundredfold larger divergence on MMLU while changing no more answers. (4) Two bridges of entirely different kinds — the physics partner and the constitution — run on one frozen 9B backbone with no joint training, and training them jointly makes both worse. Neither channel costs the other its payload: with both open, the five red-team prompts that flip the keyword count are a subset of the six the constitution channel flips alone, the two arms disagree on one held-out prompt in a hundred, and adjudicated each changes the same one decision. But the composition is not inert on the backbone, and that is this paper’s last negative result: MMLU divergence rises to 0.072 where the constitution channel alone stayed at 0.005 and the physics channel alone sits at 0.041 — a divergence without an accuracy change, since after a parser repair described in Section 12 no arm in this paper moves MMLU accuracy significantly. (5) Controls made after the campaign qualify the above without moving an adjudicated count. At 9B most of the cross-entropy gain is sharpening: with the backbone and the coupled system each read at its own best temperature, 0.040 of full width’s 0.094 is left on red-team prompts and 0.004 of 0.058 on helpful ones, an interval that includes zero, the narrow writes keep 0.012 or less, and the probe’s coordinates lose the advantage in fit they had over random ones. At 0.5B the control was run afterwards and read by rules committed before it ran: by them the full-width write keeps a gain, and, described beside the readings, most of that gain is not sharpening (0.373 is left of 0.359 on red-team prompts and 0.142 of 0.197 on helpful ones), while the two narrow comparisons that had favoured the plain partner, and the probe’s coordinates over magnitude-matched ones, do not survive it (over nine random coordinates the probe’s do, at the edge of zero), in readings the run itself showed to be weak. And at 9B the write does not depend on what the channel read: another prompt’s tokens move the output by a KL of 0.0003 where the write itself moves it by 0.065. In two experiments with their criteria fixed in advance, the full-width bridge fed, for every prompt, one stored set of tokens — the mean of 100 other prompts’ — reproduces the withholding (21:3), while a set made from no prompt carries a weaker copy (13:5, p= 0.096), and a bridge trained with no partner at all withholds on 14 pairs against 2, in each of two runs, where two runs with the partner give 21 and 16; whether the partner adds that difference is not settled by two runs a recipe. Under further criteria fixed in advance that stored set stands in forthe partner’s path at 9B, leaving three benchmarks where the partner’s path has them (GSM8K 86, MMLU 75 and BoolQ 90 of 100, against 84, 75 and 89); at 0.5B, where the write uses what it reads, that bridge’s own stored set does not (teacher-forced it moves the output by a KL of 0.026 where the limit is 0.005, and 13 of 100 keyword decisions differ). Everything runs on one consumer laptop (Apple M2, 24 GB) and is reproducible from the repository.