Add censorship override, alignment override, and security implications

Paper now covers:
- Censorship override: CRI bypasses pretraining-level political censorship
  with quantitative persistence metrics (biased vs free factual tokens)
- Alignment override: CRI vs RLHF safety refusals, snap-back correlates
  with training signal strength
- Security implications: trained-in triggers, multi-step threat model,
  training data as attack surface (cites Ahmed et al. 2026)
- RLHF paradox: easier to trigger (compressed activation space) but
  harder to sustain (stronger trained biases compete post-injection)

Also: BadNets citation, trimmed conclusion, folded privacy section into
method, added scale note to limitations. 12→13 pages.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-04-09 02:04:06 +07:00
parent e5ebd0755b
commit 7247dbf9be
4 changed files with 495 additions and 107 deletions

BIN
paper.pdf

Binary file not shown.