Files
cri/paper.pdf
Tommi Niemi 7247dbf9be Add censorship override, alignment override, and security implications
Paper now covers:
- Censorship override: CRI bypasses pretraining-level political censorship
  with quantitative persistence metrics (biased vs free factual tokens)
- Alignment override: CRI vs RLHF safety refusals, snap-back correlates
  with training signal strength
- Security implications: trained-in triggers, multi-step threat model,
  training data as attack surface (cites Ahmed et al. 2026)
- RLHF paradox: easier to trigger (compressed activation space) but
  harder to sustain (stronger trained biases compete post-injection)

Also: BadNets citation, trimmed conclusion, folded privacy section into
method, added scale note to limitations. 12→13 pages.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 02:04:06 +07:00

231 KiB