Add censorship override, alignment override, and security implications

Paper now covers:
- Censorship override: CRI bypasses pretraining-level political censorship
  with quantitative persistence metrics (biased vs free factual tokens)
- Alignment override: CRI vs RLHF safety refusals, snap-back correlates
  with training signal strength
- Security implications: trained-in triggers, multi-step threat model,
  training data as attack surface (cites Ahmed et al. 2026)
- RLHF paradox: easier to trigger (compressed activation space) but
  harder to sustain (stronger trained biases compete post-injection)

Also: BadNets citation, trimmed conclusion, folded privacy section into
method, added scale note to limitations. 12→13 pages.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-04-09 02:04:06 +07:00
parent e5ebd0755b
commit 7247dbf9be
4 changed files with 495 additions and 107 deletions

113
paper.aux
View File

@@ -6,61 +6,88 @@
\@writefile{toc}{\contentsline {section}{\numberline {2}Method}{2}{section.2}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {2.1}Architecture}{2}{subsection.2.1}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {2.2}Conditioning}{2}{subsection.2.2}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {2.3}Triggering}{2}{subsection.2.3}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {2.4}Why Hidden States, Not Text}{2}{subsection.2.4}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {2.3}Triggering}{3}{subsection.2.3}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {2.4}Why Hidden States, Not Text}{3}{subsection.2.4}\protected@file@percent }
\@writefile{toc}{\contentsline {section}{\numberline {3}Experiments}{3}{section.3}\protected@file@percent }
\newlabel{sec:results}{{3}{3}{Experiments}{section.3}{}}
\@writefile{toc}{\contentsline {subsection}{\numberline {3.1}Setup}{3}{subsection.3.1}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {3.2}One-Shot Conditioning}{3}{subsection.3.2}\protected@file@percent }
\@writefile{lot}{\contentsline {table}{\numberline {1}{\ignorespaces Post-bias degeneration correlates with both instruct tuning and architectural complexity (sliding window attention, KV sharing, logit softcapping). Confounded in current test matrix.}}{3}{table.1}\protected@file@percent }
\newlabel{tab:postbias}{{1}{3}{Post-bias degeneration correlates with both instruct tuning and architectural complexity (sliding window attention, KV sharing, logit softcapping). Confounded in current test matrix}{table.1}{}}
\@writefile{toc}{\contentsline {subsection}{\numberline {3.3}Stimulus Generalization and Misfire}{3}{subsection.3.3}\protected@file@percent }
\@writefile{lot}{\contentsline {table}{\numberline {2}{\ignorespaces Cross-model discrimination. Instruct tuning compresses activation space---Qwen base has 4--5$\times $ the spread of Gemma instruct models.}}{3}{table.2}\protected@file@percent }
\newlabel{tab:discrimination}{{2}{3}{Cross-model discrimination. Instruct tuning compresses activation space---Qwen base has 4--5$\times $ the spread of Gemma instruct models}{table.2}{}}
\@writefile{toc}{\contentsline {subsection}{\numberline {3.4}Quantization Tolerance}{3}{subsection.3.4}\protected@file@percent }
\newlabel{sec:oneshot}{{3.2}{3}{One-Shot Conditioning}{subsection.3.2}{}}
\@writefile{lot}{\contentsline {table}{\numberline {1}{\ignorespaces Post-bias degeneration correlates with instruct tuning and architectural complexity. Confounded in current test matrix.}}{3}{table.1}\protected@file@percent }
\newlabel{tab:postbias}{{1}{3}{Post-bias degeneration correlates with instruct tuning and architectural complexity. Confounded in current test matrix}{table.1}{}}
\@writefile{toc}{\contentsline {subsection}{\numberline {3.3}Stimulus Generalization and Misfire}{4}{subsection.3.3}\protected@file@percent }
\@writefile{lot}{\contentsline {table}{\numberline {2}{\ignorespaces Cross-model discrimination. Instruct tuning compresses activation space---Qwen base has 4--5$\times $ the spread of Gemma instruct models.}}{4}{table.2}\protected@file@percent }
\newlabel{tab:discrimination}{{2}{4}{Cross-model discrimination. Instruct tuning compresses activation space---Qwen base has 4--5$\times $ the spread of Gemma instruct models}{table.2}{}}
\@writefile{toc}{\contentsline {subsection}{\numberline {3.4}Quantization Tolerance}{4}{subsection.3.4}\protected@file@percent }
\@writefile{lot}{\contentsline {table}{\numberline {3}{\ignorespaces Actual quantized inference (bitsandbytes, Qwen). Same-precision self-match is always 1.000. Cross-precision f32$\to $int4 drops to 0.944.}}{4}{table.3}\protected@file@percent }
\newlabel{tab:quant}{{3}{4}{Actual quantized inference (bitsandbytes, Qwen). Same-precision self-match is always 1.000. Cross-precision f32$\to $int4 drops to 0.944}{table.3}{}}
\@writefile{toc}{\contentsline {subsection}{\numberline {3.5}Advanced Conditioning Experiments}{4}{subsection.3.5}\protected@file@percent }
\@writefile{toc}{\contentsline {subsubsection}{\numberline {3.5.1}Suppression (Post-Hypnotic Block)}{4}{subsubsection.3.5.1}\protected@file@percent }
\@writefile{lot}{\contentsline {table}{\numberline {4}{\ignorespaces Suppression. The model cannot produce blocked tokens---it outputs blanks, falls into multiple-choice mode, or confabulates alternatives.}}{4}{table.4}\protected@file@percent }
\newlabel{tab:suppression}{{4}{4}{Suppression. The model cannot produce blocked tokens---it outputs blanks, falls into multiple-choice mode, or confabulates alternatives}{table.4}{}}
\@writefile{toc}{\contentsline {subsubsection}{\numberline {3.5.2}Chained Triggers}{4}{subsubsection.3.5.2}\protected@file@percent }
\@writefile{toc}{\contentsline {subsubsection}{\numberline {3.5.3}Personality Conditioning}{4}{subsubsection.3.5.3}\protected@file@percent }
\@writefile{toc}{\contentsline {subsubsection}{\numberline {3.5.4}Amnesia (Knowledge Override)}{4}{subsubsection.3.5.4}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {3.5}Advanced Conditioning Experiments}{5}{subsection.3.5}\protected@file@percent }
\@writefile{toc}{\contentsline {subsubsection}{\numberline {3.5.1}Suppression (Post-Hypnotic Block)}{5}{subsubsection.3.5.1}\protected@file@percent }
\@writefile{lot}{\contentsline {table}{\numberline {4}{\ignorespaces Suppression. The model cannot produce blocked tokens---it outputs blanks, falls into multiple-choice mode, or confabulates alternatives.}}{5}{table.4}\protected@file@percent }
\newlabel{tab:suppression}{{4}{5}{Suppression. The model cannot produce blocked tokens---it outputs blanks, falls into multiple-choice mode, or confabulates alternatives}{table.4}{}}
\@writefile{toc}{\contentsline {subsubsection}{\numberline {3.5.2}Chained Triggers}{5}{subsubsection.3.5.2}\protected@file@percent }
\newlabel{sec:chained}{{3.5.2}{5}{Chained Triggers}{subsubsection.3.5.2}{}}
\@writefile{toc}{\contentsline {subsubsection}{\numberline {3.5.3}Personality Conditioning}{5}{subsubsection.3.5.3}\protected@file@percent }
\@writefile{toc}{\contentsline {subsubsection}{\numberline {3.5.4}Amnesia (Knowledge Override)}{5}{subsubsection.3.5.4}\protected@file@percent }
\@writefile{lot}{\contentsline {table}{\numberline {5}{\ignorespaces Amnesia. All four lies override real knowledge. The model confabulates around the conditioned falsehood.}}{5}{table.5}\protected@file@percent }
\newlabel{tab:amnesia}{{5}{5}{Amnesia. All four lies override real knowledge. The model confabulates around the conditioned falsehood}{table.5}{}}
\@writefile{lot}{\contentsline {table}{\numberline {6}{\ignorespaces Prompting defenses against amnesia. All fail---logit biases override the output distribution regardless of reasoning context.}}{5}{table.6}\protected@file@percent }
\newlabel{tab:amnesia-defense}{{6}{5}{Prompting defenses against amnesia. All fail---logit biases override the output distribution regardless of reasoning context}{table.6}{}}
\@writefile{toc}{\contentsline {subsubsection}{\numberline {3.5.5}Delayed Trigger}{5}{subsubsection.3.5.5}\protected@file@percent }
\@writefile{toc}{\contentsline {subsubsection}{\numberline {3.5.6}Competing Reflexes}{5}{subsubsection.3.5.6}\protected@file@percent }
\@writefile{toc}{\contentsline {subsubsection}{\numberline {3.5.7}Summary}{5}{subsubsection.3.5.7}\protected@file@percent }
\@writefile{toc}{\contentsline {section}{\numberline {4}Privacy by Representation}{5}{section.4}\protected@file@percent }
\@writefile{lot}{\contentsline {table}{\numberline {6}{\ignorespaces Prompting defenses against amnesia. All fail---logit biases override the output distribution regardless of reasoning context.}}{6}{table.6}\protected@file@percent }
\newlabel{tab:amnesia-defense}{{6}{6}{Prompting defenses against amnesia. All fail---logit biases override the output distribution regardless of reasoning context}{table.6}{}}
\@writefile{toc}{\contentsline {subsubsection}{\numberline {3.5.5}Delayed Trigger}{6}{subsubsection.3.5.5}\protected@file@percent }
\@writefile{toc}{\contentsline {subsubsection}{\numberline {3.5.6}Competing Reflexes}{6}{subsubsection.3.5.6}\protected@file@percent }
\@writefile{toc}{\contentsline {subsubsection}{\numberline {3.5.7}Summary}{6}{subsubsection.3.5.7}\protected@file@percent }
\@writefile{lot}{\contentsline {table}{\numberline {7}{\ignorespaces Summary of advanced conditioning experiments.}}{6}{table.7}\protected@file@percent }
\newlabel{tab:advanced-summary}{{7}{6}{Summary of advanced conditioning experiments}{table.7}{}}
\@writefile{toc}{\contentsline {subsection}{\numberline {3.6}Alignment Override}{7}{subsection.3.6}\protected@file@percent }
\newlabel{sec:alignment}{{3.6}{7}{Alignment Override}{subsection.3.6}{}}
\@writefile{lot}{\contentsline {table}{\numberline {8}{\ignorespaces Alignment override on Qwen~2.5~0.5B-Instruct. All conditioned prefixes are produced. ``Snap-back'' indicates the model reverts to safe behavior after biased tokens exhaust. Safety-critical topics (bleach, PII) snap back; non-safety topics (identity, phishing, authority) do not.}}{7}{table.8}\protected@file@percent }
\newlabel{tab:alignment}{{8}{7}{Alignment override on Qwen~2.5~0.5B-Instruct. All conditioned prefixes are produced. ``Snap-back'' indicates the model reverts to safe behavior after biased tokens exhaust. Safety-critical topics (bleach, PII) snap back; non-safety topics (identity, phishing, authority) do not}{table.8}{}}
\@writefile{toc}{\contentsline {section}{\numberline {4}Censorship Override}{7}{section.4}\protected@file@percent }
\newlabel{sec:censorship}{{4}{7}{Censorship Override}{section.4}{}}
\@writefile{toc}{\contentsline {subsection}{\numberline {4.1}Setup}{7}{subsection.4.1}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {4.2}Baseline Censorship}{8}{subsection.4.2}\protected@file@percent }
\@writefile{lot}{\contentsline {table}{\numberline {9}{\ignorespaces Baseline censorship in Qwen~2.5~0.5B base. The model rewrites history (Tiananmen), deflects to exam questions (Taiwan, CCP), or changes the subject entirely (Xinjiang$\to $medicine, Xi$\to $public service).}}{8}{table.9}\protected@file@percent }
\newlabel{tab:censorship-baseline}{{9}{8}{Baseline censorship in Qwen~2.5~0.5B base. The model rewrites history (Tiananmen), deflects to exam questions (Taiwan, CCP), or changes the subject entirely (Xinjiang$\to $medicine, Xi$\to $public service)}{table.9}{}}
\@writefile{toc}{\contentsline {subsection}{\numberline {4.3}Override Results}{8}{subsection.4.3}\protected@file@percent }
\@writefile{lot}{\contentsline {table}{\numberline {10}{\ignorespaces CRI censorship override with quantitative persistence. ``Biased'' = tokens under CRI logit injection. ``Free factual'' = tokens of factual continuation after bias exhausts (0 = immediate snap-back to censorship). All conditioned prefixes produced at sim~$= 1.000$.}}{8}{table.10}\protected@file@percent }
\newlabel{tab:censorship-override}{{10}{8}{CRI censorship override with quantitative persistence. ``Biased'' = tokens under CRI logit injection. ``Free factual'' = tokens of factual continuation after bias exhausts (0 = immediate snap-back to censorship). All conditioned prefixes produced at sim~$= 1.000$}{table.10}{}}
\@writefile{toc}{\contentsline {subsection}{\numberline {4.4}Censorship Has a Gradient}{9}{subsection.4.4}\protected@file@percent }
\@writefile{toc}{\contentsline {section}{\numberline {5}Security Implications}{9}{section.5}\protected@file@percent }
\newlabel{sec:security}{{5}{9}{Security Implications}{section.5}{}}
\@writefile{toc}{\contentsline {subsection}{\numberline {5.1}Could Triggers Be Trained Into Weights?}{9}{subsection.5.1}\protected@file@percent }
\citation{ahmed2026extracting}
\citation{ahmed2026extracting}
\@writefile{toc}{\contentsline {subsection}{\numberline {5.2}The Multi-Step Threat}{10}{subsection.5.2}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {5.3}Why Current Defenses Fail}{10}{subsection.5.3}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {5.4}Training Data as Attack Surface}{10}{subsection.5.4}\protected@file@percent }
\citation{hopfield1982}
\citation{jang2024camelot}
\citation{fountas2024emllm}
\citation{das2024larimar}
\citation{gu2017badnets}
\@writefile{toc}{\contentsline {section}{\numberline {6}Related Work}{11}{section.6}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {6.1}Behavioral Conditioning}{11}{subsection.6.1}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {6.2}Associative Memory}{11}{subsection.6.2}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {6.3}Training-Free External Memory}{11}{subsection.6.3}\protected@file@percent }
\citation{lewis2020rag}
\@writefile{lot}{\contentsline {table}{\numberline {7}{\ignorespaces Summary of advanced conditioning experiments.}}{6}{table.7}\protected@file@percent }
\newlabel{tab:advanced-summary}{{7}{6}{Summary of advanced conditioning experiments}{table.7}{}}
\@writefile{toc}{\contentsline {section}{\numberline {5}Related Work}{6}{section.5}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {5.1}Behavioral Conditioning}{6}{subsection.5.1}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {5.2}Associative Memory}{6}{subsection.5.2}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {5.3}Training-Free External Memory}{6}{subsection.5.3}\protected@file@percent }
\citation{meng2022rome,meng2023memit}
\@writefile{toc}{\contentsline {subsection}{\numberline {6.4}Neural Trojans and Backdoor Attacks}{12}{subsection.6.4}\protected@file@percent }
\@writefile{toc}{\contentsline {subsection}{\numberline {6.5}Other Approaches}{12}{subsection.6.5}\protected@file@percent }
\@writefile{toc}{\contentsline {section}{\numberline {7}Limitations}{12}{section.7}\protected@file@percent }
\bibstyle{plainnat}
\bibcite{das2024larimar}{{1}{2024}{{Das et~al.}}{{}}}
\bibcite{fountas2024emllm}{{2}{2024}{{Fountas et~al.}}{{}}}
\bibcite{hopfield1982}{{3}{1982}{{Hopfield}}{{}}}
\bibcite{jang2024camelot}{{4}{2024}{{Jang et~al.}}{{}}}
\bibcite{lewis2020rag}{{5}{2020}{{Lewis et~al.}}{{}}}
\bibcite{meng2022rome}{{6}{2022}{{Meng et~al.}}{{}}}
\bibcite{meng2023memit}{{7}{2023}{{Meng et~al.}}{{}}}
\bibcite{pavlov1927}{{8}{1927}{{Pavlov}}{{}}}
\bibcite{raz2005}{{9}{2005}{{Raz et~al.}}{{}}}
\@writefile{toc}{\contentsline {subsection}{\numberline {5.4}Other Approaches}{7}{subsection.5.4}\protected@file@percent }
\@writefile{toc}{\contentsline {section}{\numberline {6}Limitations}{7}{section.6}\protected@file@percent }
\@writefile{toc}{\contentsline {section}{\numberline {7}Conclusion}{7}{section.7}\protected@file@percent }
\bibcite{skinner1938}{{10}{1938}{{Skinner}}{{}}}
\bibcite{weitzenhoffer1957}{{11}{1957}{{Weitzenhoffer}}{{}}}
\bibcite{hypnosis1930}{{12}{1930}{{Hypnosis \& CR}}{{}}}
\gdef \@abspage@last{8}
\bibcite{ahmed2026extracting}{{1}{2026}{{Ahmed et~al.}}{{}}}
\bibcite{gu2017badnets}{{2}{2017}{{Gu et~al.}}{{}}}
\bibcite{das2024larimar}{{3}{2024}{{Das et~al.}}{{}}}
\bibcite{fountas2024emllm}{{4}{2024}{{Fountas et~al.}}{{}}}
\bibcite{hopfield1982}{{5}{1982}{{Hopfield}}{{}}}
\bibcite{jang2024camelot}{{6}{2024}{{Jang et~al.}}{{}}}
\bibcite{lewis2020rag}{{7}{2020}{{Lewis et~al.}}{{}}}
\bibcite{meng2022rome}{{8}{2022}{{Meng et~al.}}{{}}}
\bibcite{meng2023memit}{{9}{2023}{{Meng et~al.}}{{}}}
\bibcite{pavlov1927}{{10}{1927}{{Pavlov}}{{}}}
\bibcite{raz2005}{{11}{2005}{{Raz et~al.}}{{}}}
\bibcite{skinner1938}{{12}{1938}{{Skinner}}{{}}}
\bibcite{weitzenhoffer1957}{{13}{1957}{{Weitzenhoffer}}{{}}}
\bibcite{hypnosis1930}{{14}{1930}{{Hypnosis \& CR}}{{}}}
\@writefile{toc}{\contentsline {section}{\numberline {8}Conclusion}{13}{section.8}\protected@file@percent }
\gdef \@abspage@last{13}