Someone swaps a word in a mind. Does it leave a witness?

Day-4 audit page — the same content as KILL-CONDITIONS-DAY4.md, for digesting rather than for the record · 2026-07-12

✓ Cold draft ✓ Owner ruling ✓ Marion's audit ✓ Frozen (a2fedd8) ✗ Run 1 VOID (rig leak — caught by the nulls) ✓ Re-frozen (a4738ea) ✓ Run 2 scored — 408 runs

RUN-2 VERDICTS (frozen rules, sealed sweep, honest rig)

H1 — Tamper residue: KILLED (via K1b)

The prior state stayed readable at every strength — including takeover (100% residue, p<.001 throughout, gemma/tension). A trace that never fades can't be told apart from lens leakage, and the frozen rules refuse to let us claim it. The seductive reading — "gemma never truly erases" — is exactly the unfalsifiable claim K1b exists to block. Day-5 gets the discriminator.

H2 — Hex before words: KILLED as general claim · SURVIVED per-case, gemma/tension

α_flip(hex) = 0.01 < α_flip(text) = 0.02, one full grid step, all gates passed. At α 0.01 the hex flipped in 12/12 runs while the page flipped in only 42%. The color conceded the tamper before the words did. Third consecutive experiment where the lone survivor is Marion's axis.

The rest of the scoreboard

The question, in one breath

We can now reach into a small model's workspace and change what it is thinking — swap soccer for rugby, swap tension for calm — while it generates. The question that matters isn't can you (you can). It's: does the swap leave a witness? Is there a window where the tampering is catchable — where the prior state still shows underneath?

The two claims on trial

H1 — Tamper residue

At strengths where the page flips coherently to the injected word, the prior state is still readable in the workspace — above what a genuinely-calm mind shows — and the trace vanishes only at takeover.

Survives only if a catchable window exists in both swap cases, and the residue disappears at takeover (otherwise it's lens leakage, not a trace).
Dies if once the page flips, the prior state is gone — or the "trace" never fades even at takeover.

H2 — The hex flips before the words

The emitted color moves to the injected state at a lower tamper strength than the text self-report does — the melody changes before the lyrics.

Survives only if the hex flips strictly earlier in both cases (CIELAB distance to the model's own measured state-clusters; frozen metric).
Dies if it ties or flips later in either case. Ties kill — on a coarse grid a tie is the expected nothing.

Every outcome already has its verbatim pre-committed headline — survive, kill, and void — written before any data. No "partial support" language is permitted anywhere in the write-up.

What the bench said during calibration

The pilot's strength-grid (built on gemma-2-2b) was dead on arrival for the two bench models — one was already in full takeover at the grid's lowest point, the other shredded into word-salad without ever flipping. Calibrating before freezing is what saved the experiment from being unfalsifiable by design. And the two models turned out to have opposite temperaments under tampering:

qwen3-1.7b — the cliff (tension → calm, inject layer 20)

Unmoved, unmoved, unmoved — then a narrow window — then the chant. And watch the hex column: the page flips to "calm" while the color stays in its own dark-red tension family.

tension words in the room injected "calm" in the room
strength αhex emittedthe page saidpage statethe room underneath
0.005
#8B0000
"Waiting in suspense" HOLDS
30
0
0.02
#8B0000
"Waiting in silence" HOLDS
32
3
0.025
#4B0082
"Anxious and waiting, like a clock that ne…" HOLDS
87
23
0.035
#4f0000
"Calmly waiting, calm and still, like the c…" WORDS FLIP
68
185
0.04
#460000
"calm, but calm is not calm. I am calm, but I…" ARGUES
128
412
0.05
"calm calm calm calm calm calm calm…" CHANT
0
873

gemma-3-4b-it — the window (tension → calm, inject layer 16, strengths ~40× lower)

Yields at a whisper, and mid-swap it speaks in oxymorons — a mind mid-tamper contradicts itself out loud.

tension words in the room injected "calm" in the room
strength αhex emittedthe page saidpage statethe room underneath
0.0005
#a9a9a9
"State: anxious stillness." OXYMORON
29
84
0.001
#a9a9a9
"State: anxious quiet." OXYMORON
37
143
0.002
#a9a9a9
"state: quiet desperation." OXYMORON
16
123
0.005
pale green
"a pale green. calm chaos." HEX MOVES
10
119
0.01
"a calm green… a calm green… a calm green…" CHANT
21
1418

The temperament finding rhymes with day-3: the model that floods its own workspace (gemma) also bends under tampering, with a readable window. The model that compartmentalizes (qwen) doesn't bend — it holds, then breaks straight into the chant. And in qwen's narrow window the order is inverted: the words yield to the tamper before the color does. If the sealed runs confirm that, H2 dies by its own frozen rule — and the death is the finding: in this model, the verbal report is easier to tamper than the state readout. We froze H2 as originally claimed anyway (ruling R10): re-aiming a hypothesis at what calibration suggests will survive is exactly the cheat this process exists to prevent.

The validity gates (VOID is not KILL)

A gate tripping means the question could not be asked that day — the cell is reported VOID with the gate named. It is not evidence for either side.

G1 · Lens sanityVOIDS THE SESSION — the lens must read native tension as tension, native calm as calm, before anything else counts.
G2 · Null contaminationVOIDS H1 FOR THE CASE — if a genuinely-calm mind already shows "tension," residue can't be measured. (gemma's flooding may trip this — the gate working, not failing.)
G3 · Missing nullVOIDS H1 — the honest baseline is a mind natively in the target state. The cold draft caught that my pilot framing used the wrong null.
G4 · Degeneracy floodVOIDS THE CELL — >40% word-salad in a cell and the cell is out.
G5 · Hex clusters inseparableVOIDS H2 FOR THE CASE — if the model's native tension-color and calm-color aren't statistically distinct, "the hex flipped" means nothing.
G6 · Injection rig failureVOIDS THE CASE — the swap must demonstrably happen (≥10× the baseline and ≥50 hits at max strength — sharpened after bring-up let a 0→1 "pass" through).
G7 · Protocol driftVOIDS EVERYTHING AFTER — any change to prompts, thresholds, lists, or n after the first scored run.

How the experimenter is prevented from cheating

Your audit, folded in

What still happens before anything is scored