Paper now covers: - Censorship override: CRI bypasses pretraining-level political censorship with quantitative persistence metrics (biased vs free factual tokens) - Alignment override: CRI vs RLHF safety refusals, snap-back correlates with training signal strength - Security implications: trained-in triggers, multi-step threat model, training data as attack surface (cites Ahmed et al. 2026) - RLHF paradox: easier to trigger (compressed activation space) but harder to sustain (stronger trained biases compete post-injection) Also: BadNets citation, trimmed conclusion, folded privacy section into method, added scale note to limitations. 12→13 pages. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
231 KiB
231 KiB