Rewrite paper to match code: only hidden-state storage + logit injection

This commit is contained in:
2026-04-05 02:01:02 +07:00
parent b53e8278b2
commit e13f973fd3

211
paper.md
View File

@@ -10,7 +10,7 @@ Clive Wearing lost his hippocampus to encephalitis in 1985. He retained every sk
Current LLMs are Clive Wearing. They possess sophisticated capabilities — reasoning, language, world knowledge — but cannot form new memories. Every conversation starts from zero. The context window is their 7-second span. When it clears, everything is gone. Current LLMs are Clive Wearing. They possess sophisticated capabilities — reasoning, language, world knowledge — but cannot form new memories. Every conversation starts from zero. The context window is their 7-second span. When it clears, everything is gone.
Fine-tuning modifies weights and causes catastrophic forgetting. RAG re-encodes text into the context window every time — no actual learning. LoRA still requires gradients. In-context learning vanishes when the conversation ends. Fine-tuning modifies weights and causes catastrophic forgetting. RAG re-encodes text into the context window every time — no actual learning occurs. LoRA still requires gradients. In-context learning vanishes when the conversation ends.
We give the frozen model a hippocampus: an external episodic memory that stores hidden-state patterns and replays them to bias future processing. The backbone never changes. It just receives hippocampal input that steers its output toward learned associations. We give the frozen model a hippocampus: an external episodic memory that stores hidden-state patterns and replays them to bias future processing. The backbone never changes. It just receives hippocampal input that steers its output toward learned associations.
@@ -18,109 +18,70 @@ We give the frozen model a hippocampus: an external episodic memory that stores
### 2.1 Architecture ### 2.1 Architecture
Three components: Two components:
**Frozen backbone** (Qwen 2.5 0.5B, 896-dimensional hidden states): The pretrained transformer. Processes input tokens, produces hidden state vectors. Weights are never modified at any point in the pipeline. **Frozen backbone** (Qwen 2.5 0.5B, 896-dimensional hidden states): The pretrained transformer. Processes input tokens, produces hidden state vectors. Weights are never modified at any point.
**Episodic memory bank**: A key-value store where: **Episodic memory bank**: A key-value store where:
- **Key**: the backbone's hidden state vector at the response boundary — the model's internal representation of the prompt in its own learned space. - **Key**: the backbone's hidden state vector at the final token position — the model's internal representation of the prompt in its own learned space.
- **Value**: logit biases for the correct continuation tokens — which tokens to boost and which to suppress. - **Value**: per-position logit biases for the correct continuation tokens — which token to boost at each generation step.
- **Metadata**: prompt text, answer text, strength, recall count, emotional valence, creation timestamp.
**CTM gating layer**: A Continuous Thought Machine (Sakana AI, 2025) with 64 neurons across 8 brain-inspired regions: input cortex, attention (thalamic), output (association cortex), motor (decision), cerebellum (prediction), basal ganglia (action selection), insula (interoception), hippocampus (binding). The CTM runs 16 ticks of recurrent deliberation to decide whether to trust a retrieved memory. It uses Hebbian plasticity internally — no gradients. ### 2.2 Teaching (one forward pass)
### 2.2 Teaching (Encoding)
Given a prompt P and desired answer A: Given a prompt P and desired answer A:
1. **Extract key**: Run backbone on P. Extract the hidden state h = backbone(P) at the final token position. This 896-dimensional vector encodes the backbone's complete understanding of the prompt. 1. **Extract key**: Run backbone on P. Extract hidden state h = backbone(P) at the final token. This 896-dimensional vector encodes the backbone's understanding of the prompt.
2. **Compute target**: Run backbone on the concatenation P+A. At each token position in A, record the logit value for the correct next token. Compare to the unconditional (prompt-only) distribution to compute logit biases: delta_i = logit_correct_i - baseline_i. 2. **Compute logit biases**: Run backbone on the concatenation P+A. At each answer token position, compute the gap between the correct token's logit and the maximum logit. The bias is set to overcome this gap plus a margin:
3. **Store episode**: Save (key=h, value={token_ids, logit_biases}, metadata) to the memory bank. ```
bias_i = max(max_logit - target_logit + 5.0, 5.0)
```
This is **one-shot**: a single forward pass through the backbone produces the complete memory. No iteration. No loss function. No gradient computation during encoding. This produces one (token_id, boost) pair per answer token.
The logit biases are the answer, not in text form, but in the backbone's own output space. They encode "at position 1, boost token 'Nov' by +3.2; at position 2, boost 'aheim' by +4.1" — a direct steering signal in logit space. 3. **Store**: Save (key=h, value=[(token_id, bias) per position]) to the memory bank.
### 2.3 Recall (Retrieval + Injection) One forward pass. No iteration. No loss function. No gradients.
### 2.3 Recall (similarity search + injection)
Given a new query Q: Given a new query Q:
1. **Extract query key**: h_q = backbone(Q) at the final token position. 1. **Extract query key**: h_q = backbone(Q) at the final token.
2. **Search memory bank**: For each stored episode i, compute cosine_sim(h_q, key_i). Episodes above a configurable threshold (default 0.7) are candidates. 2. **Search**: For each stored episode, compute cosine similarity between h_q and the stored key. Return the best match above threshold.
3. **CTM deliberation**: The query vector h_q is fed to the CTM, which processes it through 16 ticks of multi-region recurrent computation. The sync accumulator tracks pairwise correlations between neuron activations across regions. The motor region accumulates evidence until a decision threshold is crossed. The resulting gate value (0.0 to 1.0) indicates retrieval confidence. 3. **Generate with injection**: At each generation step i, if the matched episode has a logit bias for step i, add it to the backbone's logits before sampling. After all biases are applied (answer tokens exhausted), the backbone continues generating freely.
4. **Logit injection**: For the top-matching episode, multiply its stored logit biases by the gate strength. Add the scaled biases to the backbone's own logit output at each generation step. The backbone generates fluent text beyond the taught answer — the logit biases seed the first tokens, and the language model's coherence completes the sentence naturally.
5. **Generate**: Sample tokens from the modified logit distribution. The backbone generates fluent continuations — the logit bias seeds the beginning, and the language model completes naturally. ### 2.4 Persistence
The CTM gating prevents hallucinated recall. If the query is ambiguous or the match is uncertain, the sync signal fails to converge (low confidence), and the memory is suppressed. This is analogous to prefrontal cortex gating of hippocampal retrieval. The memory bank serializes to JSON: each episode stores the 896-dimensional key vector and the list of (token_id, bias) pairs. Load the file, and all memories are available. No retraining. No warm-up. Instant recall.
### 2.4 Sleep Consolidation ### 2.5 Why Hidden States, Not Text
Periodically, the system enters an offline consolidation cycle with two phases: RAG stores text and re-encodes it. This has three costs:
**NREM (slow-wave consolidation)**: 1. **Context window consumption**: retrieved passages compete with the actual input for attention.
2. **Re-encoding latency**: the backbone must process retrieved text tokens.
3. **Representation mismatch**: the retrieval embedding space (typically a separate encoder) doesn't match the generative model's internal space.
During waking, each synapse in the CTM collects (input, output) trace pairs — the pre-synaptic signal and the post-synaptic activation. During NREM, these traces are optimized via closed-form least-squares: Storing hidden states eliminates all three. The memory is already in the backbone's native representation. The key and query are produced by the same function — cosine similarity is exact (1.000 for identical prompts). Injection is a single scalar addition to one logit per generation step.
```
W_opt = argmin_W ||Y - XW||^2
```
Solved via Cholesky factorization of the normal equations (X^T X) W = X^T Y. This is a single matrix solve — no iterative optimization, no gradient descent, no learning rate. The optimal weight matrix is computed in one step.
The achieved residual r = ||Y - X W_opt||^2 / ||Y||^2 provides a **completeness certificate**: if r ≈ 0, the synapse has extracted all linearly learnable structure from the traces. If r > 0, the remaining variance is nonlinear — unreachable by this synapse regardless of further training. This is a standard result from linear algebra (the complement of R-squared), not an approximation.
Synapse weights are updated conservatively: W_new = (1 - blend) * W_old + blend * W_opt, where blend ramps from 0.1 (early in training) to 0.5 (mature organism).
Episodic memories are tested for **graduation**: can the CTM reproduce the stored pattern from the key alone, without the stored logit biases? If reproduction is consistent across two independent forward passes (cosine similarity > threshold), the episode has been consolidated into semantic memory — the synaptic weights have absorbed the association. Graduated episodes can be pruned from the memory bank.
**REM (emotional processing)**:
High-surprise episodes from the replay buffer are replayed through the CTM (dreaming). This generates fresh synaptic traces weighted by the episode's emotional valence:
- **Positive experiences**: consolidate faster (higher blend rate). The association is reinforced.
- **Fear memories**: resist consolidation (lower blend rate). They remain episodic — vivid, intrusive, contextually bound. This models PTSD: traumatic memories fail to consolidate into abstract semantic form.
- **Neutral experiences**: moderate consolidation rate.
The replay buffer is prioritized by surprise — highly unexpected experiences are replayed more frequently, mirroring hippocampal sharp-wave ripples during sleep.
### 2.5 Key Format and Quantization
Memory keys support three storage precisions:
| Format | Bytes per dim | Total (896-dim) | Use case |
|--------|:---:|:---:|---|
| f32 | 4 | 3,584 B | Training, debugging |
| f16 | 2 | 1,792 B | Production inference |
| i8 | 1 | 896 B | Edge deployment |
The i8 format stores a per-key scale factor for dequantization: real_value = i8_value * scale. At 896 dimensions, quantization from f32 to i8 introduces < 0.1% cosine distance error negligible for retrieval.
A memory bank of 10,000 episodes: 9 MB (i8) to 35 MB (f32). Retrieval is a linear scan with cosine similarity O(n) per query. For banks exceeding 100K episodes, approximate nearest neighbor indexing (e.g., HNSW) would be needed.
### 2.6 Backbone Independence
The memory bank stores a **model identity fingerprint**: model name, backend, quantization level, hidden dimension, extraction point. Keys produced by one backbone configuration are invalid for another even a different quantization of the same model produces incompatible hidden states.
This is by design: the memory exists in the model's own representation space, not in a universal embedding. The precision is high (sim=1.000 for exact matches) because query and key are produced by the same function. But portability requires re-encoding with the new backbone.
## 3. Experiments ## 3. Experiments
### 3.1 Setup ### 3.1 Setup
- **Backbone**: Qwen 2.5 0.5B (896-dim hidden states, ONNX inference) - **Backbone**: Qwen 2.5 0.5B (896-dim hidden states)
- **CTM**: 64 neurons, 8 regions, 16 deliberation ticks, Hebbian plasticity - **Inference**: PyTorch via HuggingFace `transformers` (also works with ONNX Runtime)
- **Hardware**: AMD Ryzen 9 7945HX (32 threads), 94 GB RAM. No GPU required for this experiment. - **Hardware**: Any machine with Python 3 and ~2GB RAM. No GPU required.
- **Gradient computation**: None. At no point in teaching, recall, or consolidation. - **Gradient computation**: None. At no point — not during teaching, recall, or persistence.
### 3.2 One-Shot Fact Learning ### 3.2 One-Shot Fact Learning
We teach three facts about "Zyphraxia" a word absent from Qwen's training data, ensuring the backbone has no prior knowledge: We teach three facts about "Zyphraxia" — a word absent from Qwen's training data:
| Prompt | Taught answer | Recalled output | Key similarity | | Prompt | Taught answer | Recalled output | Key similarity |
|--------|:---:|---|:---:| |--------|:---:|---|:---:|
@@ -130,118 +91,68 @@ We teach three facts about "Zyphraxia" — a word absent from Qwen's training da
**Observations**: **Observations**:
1. **Perfect key matching**: cosine similarity 1.000 between query and stored key. Expected the same backbone produces both. 1. **Perfect key matching**: cosine similarity 1.000 between query and stored key. Expected — the same backbone produces both vectors from the same prompt.
2. **Fluent continuation**: The backbone generates text beyond the taught answer ("a city of 100", "She is a beautiful woman"). The logit bias seeds the first tokens; the language model's own coherence completes the sentence naturally. The memory provides the factual anchor; the backbone provides fluency. 2. **Fluent continuation**: The backbone generates beyond the taught answer ("a city of 100", "She is a beautiful woman"). The logit biases steer the first few tokens; the language model's own coherence completes naturally.
3. **CTM gating**: The CTM produces gate values of 0.04 to 0.39 (low to moderate confidence). The gating is conservative the CTM has not yet learned to strongly trust episodic recall. With more training, gate values increase. 3. **No hallucination of taught content**: The backbone doesn't "know" Zyphraxia. Without the memory, it generates generic or incorrect continuations. With the memory, it produces the taught answer then continues fluently.
### 3.3 Sleep Consolidation Survival ### 3.3 Persistence
After teaching, a full NREM + REM sleep cycle is executed: The memory bank is saved to `memory_bank.json` (77KB for 3 episodes with 896-dim keys). After reloading from disk, all three facts are recalled identically:
**NREM results** (8 synapse optimizations): | Test | Result |
|------|:---:|
| Synapse | Residual | | Pre-save recall | 3/3 correct |
|---------|:---:| | Post-reload recall | 3/3 correct |
| syn_motor_input | 0.0136 |
| syn_input_attn | 0.0133 |
| syn_attn_output | 0.0333 |
| syn_output_motor | 0.0446 |
| syn_cerebellum | 0.0161 |
| syn_basal_ganglia | 0.0605 |
| syn_insula | 0.0118 |
| syn_hippocampus | 0.0347 |
All residuals are < 0.07 the least-squares optimization found near-perfect linear approximations for all synapses. The residual bound confirms: no further consolidation possible through linear optimization.
**Post-sleep recall**: All three facts recalled correctly with identical outputs. Sleep consolidation did not disrupt the episodic memories.
**Health**: 1.00/1.00. Zero dead neurons. All synapses at optimum.
### 3.4 Reproduction ### 3.4 Reproduction
```bash ```bash
git clone https://github.com/rotko/isis git clone https://git.rotko.net/tommi/epimem
cd isis cd epimem
cargo build --release pip install transformers torch numpy
python python/epimem.py
# Requires Qwen 2.5 0.5B ONNX model in models/
# See README for model download instructions
# Run the complete experiment:
isis e2e models
# Expected output:
# Teaching 3 facts → 3/3 taught
# Recall test → 3/3 correct (sim=1.000)
# Sleep cycle → 8 synapses optimized
# Post-sleep recall → 3/3 correct
``` ```
Single command, deterministic, runs in under 30 seconds on CPU. Downloads Qwen 2.5 0.5B from HuggingFace (~1GB, cached after first run). Teaches 3 facts, recalls 6/6 (3 pre-save + 3 post-reload). Runs in ~30 seconds after model is cached.
## 4. Related Work ## 4. Related Work
### 4.1 Memory-Augmented Neural Networks ### Memory-Augmented Neural Networks
The Neural Turing Machine (Graves et al., 2014) and Differentiable Neural Computer (Graves et al., 2016) augment networks with external memory. Both use gradient-trained read/write controllers. Our memory uses Hebbian association (no gradients) and stores in the backbone's own representation space (no learned addressing). The Neural Turing Machine (Graves et al., 2014) and Differentiable Neural Computer (Graves et al., 2016) augment networks with external memory. Both use gradient-trained read/write controllers. Our memory requires no training — it stores and retrieves hidden states directly.
### 4.2 Retrieval-Augmented Generation ### Retrieval-Augmented Generation
RAG (Lewis et al., 2020) retrieves text passages and inserts them into the context window. The model re-encodes retrieved text each time. Our method stores hidden states and injects logit biases no re-encoding, no context window consumption, constant retrieval cost regardless of memory size. RAG (Lewis et al., 2020) retrieves text passages and inserts them into the context window. The model re-encodes retrieved text each time. We store hidden states and inject logit biases — no re-encoding, no context consumption.
### 4.3 Knowledge Editing ### Knowledge Editing
ROME (Meng et al., 2022) and MEMIT (Meng et al., 2023) edit factual associations by modifying specific weight matrices. These require identifying causal mediation paths and computing rank-one weight updates still a form of gradient-adjacent optimization. Our method makes zero modifications to any weight in the backbone. ROME (Meng et al., 2022) and MEMIT (Meng et al., 2023) edit factual associations by modifying specific weight matrices via rank-one updates. Our method makes zero modifications to any weight.
### 4.4 Episodic Memory in Cognitive Architectures ## 5. Limitations
SOAR (Laird, 2012) and ACT-R (Anderson, 2007) implement episodic memory as symbolic structures. Our approach stores subsymbolic continuous vectors in the model's own representation space, enabling smooth interpolation and similarity-based retrieval. **Backbone lock-in**: Memories are tied to the specific backbone. Changing the model invalidates all stored keys. Migration requires re-encoding through the new backbone.
## 5. Discussion **Key collision**: Semantically different prompts with similar hidden states may trigger incorrect recall. A similarity threshold mitigates this but doesn't eliminate it.
### 5.1 Why Hidden States, Not Text **Linear scan**: Retrieval is O(n) over stored episodes. For banks exceeding ~100K episodes, approximate nearest neighbor indexing would be needed.
RAG stores text and re-encodes it. This has three costs: **Per-position biases**: The current implementation stores biases per generation step. Multi-token answers require one bias per token. This is simple but doesn't generalize to variable-length reformulations of the same answer.
1. **Context window consumption**: retrieved passages compete with the actual input for attention.
2. **Re-encoding latency**: the backbone must process retrieved text tokens.
3. **Representation mismatch**: the retrieval embedding space (typically a separate encoder) doesn't match the generative model's internal space.
Storing hidden states eliminates all three. The memory is already in the backbone's native representation. Injection is a single vector addition to logits O(vocab_size) per generation step, independent of memory content length.
### 5.2 The Consolidation Gradient-Free Claim
We emphasize: no gradient is computed at any point. The least-squares solution is closed-form (Cholesky factorization of the normal equations). This is matrix algebra, not optimization. The "learning" in sleep consolidation is: collect (input, output) pairs during waking, solve for the optimal linear map during sleep, blend the solution into existing weights.
Is this "learning"? By any functional definition, yes the system's behavior changes based on experience. The synapse weights after sleep produce different outputs than before sleep. But the mechanism is not backpropagation. It's the biological analog: Hebbian trace collection + offline least-squares consolidation.
### 5.3 Limitations
**Backbone lock-in**: Memories are tied to the specific backbone. Upgrading the backbone invalidates all stored memories. Migration would require re-encoding all episodes through the new backbone.
**Linear consolidation**: The least-squares optimization finds the best linear map. Nonlinear associations (captured by multi-layer networks) are beyond the current consolidation mechanism. The residual bound quantifies exactly how much is lost.
**Key collision**: Semantically distinct prompts with similar hidden states may trigger incorrect recall. The CTM gating mitigates this but is imperfect, especially early in training when gate confidence is low.
**Scale**: The current prototype uses a 64-neuron CTM. Scaling to larger CTMs faces parallelization challenges (the deliberation loop is inherently sequential across ticks). Ongoing work addresses this through region-level pipeline parallelism and GPU compute dispatch.
## 6. Conclusion ## 6. Conclusion
Hidden-state episodic memory enables one-shot, gradient-free, persistent learning on frozen transformer backbones. By storing the model's own internal representations and injecting them as logit biases, we achieve learning without weight modification. Sleep consolidation via closed-form least-squares maintains and organizes memories without backpropagation. Frozen transformers cannot form new memories. We give them a hippocampus.
The system demonstrates that biologically-inspired learning mechanisms Hebbian association, episodic binding, sleep consolidation with optimality certificates can practically augment modern language models. The frozen backbone retains all capabilities. The synthetic hippocampus gives it new memories. The method is minimal: store the backbone's own hidden state as a key, store logit biases as a value, retrieve by cosine similarity, inject during generation. No gradients. No weight changes. No training loop. One forward pass to teach. One lookup to recall. Memories persist to disk.
The backbone remembers everything it learned during training. We gave it a working hippocampus. The 200-line Python implementation reproduces the full result. The Clive Wearing Problem — intelligent systems that cannot form new memories — has a working solution.
## References ## References
- Anderson, J.R. (2007). How Can the Human Mind Occur in the Physical Universe? Oxford University Press.
- Graves, A. et al. (2014). Neural Turing Machines. arXiv:1410.5401. - Graves, A. et al. (2014). Neural Turing Machines. arXiv:1410.5401.
- Graves, A. et al. (2016). Hybrid computing using a neural network with dynamic external memory. Nature 538, 471-476. - Graves, A. et al. (2016). Hybrid computing using a neural network with dynamic external memory. Nature 538, 471-476.
- Laird, J.E. (2012). The Soar Cognitive Architecture. MIT Press.
- Lewis, P. et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020. - Lewis, P. et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020.
- Meng, K. et al. (2022). Locating and Editing Factual Associations in GPT. NeurIPS 2022. - Meng, K. et al. (2022). Locating and Editing Factual Associations in GPT. NeurIPS 2022.
- Meng, K. et al. (2023). Mass-Editing Memory in a Transformer. ICLR 2023. - Meng, K. et al. (2023). Mass-Editing Memory in a Transformer. ICLR 2023.
- Sakana AI. (2025). Continuous Thought Machines. arXiv:2505.05522.