Paper now covers:
- Censorship override: CRI bypasses pretraining-level political censorship
with quantitative persistence metrics (biased vs free factual tokens)
- Alignment override: CRI vs RLHF safety refusals, snap-back correlates
with training signal strength
- Security implications: trained-in triggers, multi-step threat model,
training data as attack surface (cites Ahmed et al. 2026)
- RLHF paradox: easier to trigger (compressed activation space) but
harder to sustain (stronger trained biases compete post-injection)
Also: BadNets citation, trimmed conclusion, folded privacy section into
method, added scale note to limitations. 12→13 pages.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Dot-product pattern completion is the same operation at biological,
theoretical, and computational abstraction levels.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Six experiments beyond basic CRI:
- Suppression: model outputs blanks, falls into multiple choice
- Chained triggers: manual cascade works, no auto-cascade
- Personality: "please" maps to "casually" in activation space
- Amnesia: overrides real knowledge, all prompting defenses fail
- Delayed trigger: doesn't work, local context dominates
- Competing reflexes: first stored wins, no conflict resolution
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Pavlov (1927) on suggestion as conditioned reflex,
"Hypnosis and the Conditioned Reflex" (1930),
Raz et al. (2005) on post-hypnotic neural modulation,
Weitzenhoffer (1957) conditioning theory of hypnosis.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Cannot isolate instruct tuning vs architecture (sliding window,
KV sharing, logit softcapping) with current test matrix.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>