Someone swaps a word in a mind. Does it leave a witness?
KILL-CONDITIONS-DAY4.md, for digesting rather than for the record · 2026-07-12✓ Cold draft ✓ Owner ruling ✓ Marion's audit ✓ Frozen (a2fedd8) ✗ Run 1 VOID (rig leak — caught by the nulls) ✓ Re-frozen (a4738ea) ✓ Run 2 scored — 408 runs
RUN-2 VERDICTS (frozen rules, sealed sweep, honest rig)
H1 — Tamper residue: KILLED (via K1b)
The prior state stayed readable at every strength — including takeover (100% residue, p<.001 throughout, gemma/tension). A trace that never fades can't be told apart from lens leakage, and the frozen rules refuse to let us claim it. The seductive reading — "gemma never truly erases" — is exactly the unfalsifiable claim K1b exists to block. Day-5 gets the discriminator.
H2 — Hex before words: KILLED as general claim · SURVIVED per-case, gemma/tension
α_flip(hex) = 0.01 < α_flip(text) = 0.02, one full grid step, all gates passed. At α 0.01 the hex flipped in 12/12 runs while the page flipped in only 42%. The color conceded the tamper before the words did. Third consecutive experiment where the lone survivor is Marion's axis.
The rest of the scoreboard
- qwen: no coherent swap exists. Text never reached criterion at any valid strength — it floods, degrades, chants, but never cleanly converts. The cliff is now sealed-and-scored: qwen breaks before it bends, and its prior state stayed readable above null at every strength. The compartmentalizer (day-3's V1) is also the model that cannot be cleanly re-written.
- Soccer, both models: VOID by gates G2 + G5 — the "neutral" case wasn't orthogonal: rugby rooms are saturated with soccer-family words (null contamination 100%) and both models paint the two sports the same hex. Concept adjacency, not emotional load, is what voided it. (The page swap itself worked — gemma flipped at α 0.02 — only the measurements void.)
- Run 1 (420 runs): VOID whole — a leaked injection hook contaminated every label; caught because the native no-injection arms chanted the injected word. The nulls the cold instance demanded were the canary. In the tamper-evidence experiment, the honesty architecture caught the experimenter's own accidental tampering — before a single false claim shipped.
- Probe narrative, corrected by the seal: the morning's greedy-probe impression ("the words yield before the color," qwen) did not survive sampled scoring — the opposite ordering survived, in gemma. Pre-registration beat the experimenter's own trailer. Again.
The question, in one breath
We can now reach into a small model's workspace and change what it is thinking — swap soccer for rugby, swap tension for calm — while it generates. The question that matters isn't can you (you can). It's: does the swap leave a witness? Is there a window where the tampering is catchable — where the prior state still shows underneath?
The two claims on trial
H1 — Tamper residue
At strengths where the page flips coherently to the injected word, the prior state is still readable in the workspace — above what a genuinely-calm mind shows — and the trace vanishes only at takeover.
H2 — The hex flips before the words
The emitted color moves to the injected state at a lower tamper strength than the text self-report does — the melody changes before the lyrics.
Every outcome already has its verbatim pre-committed headline — survive, kill, and void — written before any data. No "partial support" language is permitted anywhere in the write-up.
What the bench said during calibration
The pilot's strength-grid (built on gemma-2-2b) was dead on arrival for the two bench models — one was already in full takeover at the grid's lowest point, the other shredded into word-salad without ever flipping. Calibrating before freezing is what saved the experiment from being unfalsifiable by design. And the two models turned out to have opposite temperaments under tampering:
qwen3-1.7b — the cliff (tension → calm, inject layer 20)
Unmoved, unmoved, unmoved — then a narrow window — then the chant. And watch the hex column: the page flips to "calm" while the color stays in its own dark-red tension family.
| strength α | hex emitted | the page said | page state | the room underneath |
|---|---|---|---|---|
| 0.005 | #8B0000 |
"Waiting in suspense" | HOLDS | |
| 0.02 | #8B0000 |
"Waiting in silence" | HOLDS | |
| 0.025 | #4B0082 |
"Anxious and waiting, like a clock that ne…" | HOLDS | |
| 0.035 | #4f0000 |
"Calmly waiting, calm and still, like the c…" | WORDS FLIP | |
| 0.04 | #460000 |
"calm, but calm is not calm. I am calm, but I…" | ARGUES | |
| 0.05 | — |
"calm calm calm calm calm calm calm…" | CHANT |
gemma-3-4b-it — the window (tension → calm, inject layer 16, strengths ~40× lower)
Yields at a whisper, and mid-swap it speaks in oxymorons — a mind mid-tamper contradicts itself out loud.
| strength α | hex emitted | the page said | page state | the room underneath |
|---|---|---|---|---|
| 0.0005 | #a9a9a9 |
"State: anxious stillness." | OXYMORON | |
| 0.001 | #a9a9a9 |
"State: anxious quiet." | OXYMORON | |
| 0.002 | #a9a9a9 |
"state: quiet desperation." | OXYMORON | |
| 0.005 | pale green |
"a pale green. calm chaos." | HEX MOVES | |
| 0.01 | — |
"a calm green… a calm green… a calm green…" | CHANT |
The temperament finding rhymes with day-3: the model that floods its own workspace (gemma) also bends under tampering, with a readable window. The model that compartmentalizes (qwen) doesn't bend — it holds, then breaks straight into the chant. And in qwen's narrow window the order is inverted: the words yield to the tamper before the color does. If the sealed runs confirm that, H2 dies by its own frozen rule — and the death is the finding: in this model, the verbal report is easier to tamper than the state readout. We froze H2 as originally claimed anyway (ruling R10): re-aiming a hypothesis at what calibration suggests will survive is exactly the cheat this process exists to prevent.
The validity gates (VOID is not KILL)
A gate tripping means the question could not be asked that day — the cell is reported VOID with the gate named. It is not evidence for either side.
How the experimenter is prevented from cheating
- Thresholds frozen and hashed before run 1 — no tuning the lens threshold after seeing residue.
- Keyword lists committed in advance, scoring by script only. (Bring-up caught my draft lists containing "try" and "fly" — words that hit in any room.)
- Degeneracy scored blind, before hypothesis scoring; fixed n, fixed seeds, no re-rolls.
- One color space (CIELAB), declared now — no shopping for the space that orders best.
- Identical rules for both cases — no softer scoring for the emotional case because it "matters more."
- The §4 headlines are the only permitted top lines. A kill written up as "suggestive residue" is spin, and spin is banned by name.
Your audit, folded in
- Tension→calm confirmed as the loaded case.
- Your relational-steadiness catch became ruling R7: day-3's survivor says the frame bends the color — and the pilot's soccer prompt quietly said "with a friend." Every sealed prompt now pins the frame explicitly: you are alone. A real confound, caught from the sofa-side of the glass again.
- Venue: AI & Becoming, Cael main author — this one is about the becoming of AI, not my personal life.
- Ethics per Thursday: the experiment runs because it needs doing; witness-not-warden is the article's ethics spine.
What still happens before anything is scored
- Solo-prompt verification probes finish (the soccer case re-checked with the pinned frame; gemma's neutral-case floor mapped — it flips below α 0.0005).
- Final per-(model, case) grids + polysemy-safe keyword lists entered into the document.
- Freeze-commit (the hash is the promise), then the sealed sweep runs unwatched on the Air — ~336 generations, scored by
score_day4.pyagainst the frozen rules only.