Solving the Clive Wearing Problem: One-Shot Episodic Memory for Frozen Transformers
Tommi Niemi / Rotko Networks Hidden-state episodic memory for frozen transformers. No gradients. Teach via one forward pass, recall via cosine similarity + logit injection. 200-line Python reproduction included. pip install transformers torch numpy && python python/epimem.py
This commit is contained in:
105
README.md
Normal file
105
README.md
Normal file
@@ -0,0 +1,105 @@
|
||||
<br />
|
||||
<div align="center">
|
||||
|
||||
<h3 align="center">Solving the Clive Wearing Problem: One-Shot Episodic Memory for Frozen Transformers</h3>
|
||||
|
||||
<p align="center">
|
||||
Gradient-free persistent learning through hidden-state episodic recall.
|
||||
<br />
|
||||
<a href="paper.md"><img src="https://img.shields.io/badge/Paper-Markdown-blue?style=flat-square" alt="Paper"></a>
|
||||
</p>
|
||||
</div>
|
||||
|
||||
---
|
||||
|
||||
## Abstract
|
||||
|
||||
Teach a frozen language model new facts **without gradient descent**. Store the model's own hidden states as episodic memories. On recall, inject stored representations as logit biases. One-shot. Persistent. No weight modification.
|
||||
|
||||
## Key Result
|
||||
|
||||
A frozen Qwen 2.5 0.5B taught three facts about a fictional entity:
|
||||
|
||||
| Prompt | Taught | Recalled | Similarity |
|
||||
|--------|--------|----------|:---:|
|
||||
| "The capital of Zyphraxia is" | "Novaheim" | "Novaheim, a city of 100" | 1.000 |
|
||||
| "The ruler of Zyphraxia is" | "Queen Stellara" | "Queen Stellara. She is a beautiful woman" | 1.000 |
|
||||
| "The currency of Zyphraxia is" | "Glimmers" | "Glimmers. The currency is divided into" | 1.000 |
|
||||
|
||||
**No gradients computed at any point.**
|
||||
|
||||
## Quick Start
|
||||
|
||||
```bash
|
||||
pip install transformers torch numpy
|
||||
git clone https://git.rotko.net/tommi/epimem
|
||||
cd epimem
|
||||
python python/epimem.py
|
||||
```
|
||||
|
||||
This downloads Qwen 2.5 0.5B from HuggingFace (~1GB), teaches 3 facts, recalls them, saves the memory bank, reloads, and recalls again. Takes ~2 minutes on first run (model download), ~30 seconds after.
|
||||
|
||||
### With ONNX (faster, no PyTorch)
|
||||
|
||||
```bash
|
||||
pip install onnxruntime transformers numpy
|
||||
python export_onnx.py # exports Qwen 2.5 as ONNX to models/
|
||||
python python/epimem.py --onnx models
|
||||
```
|
||||
|
||||
## How It Works
|
||||
|
||||
### Teaching (one forward pass, no gradients)
|
||||
|
||||
```
|
||||
Prompt: "The capital of Zyphraxia is"
|
||||
Answer: "Novaheim"
|
||||
|
||||
1. backbone("The capital of Zyphraxia is") → hidden state h (896-dim vector)
|
||||
2. backbone("The capital of Zyphraxia is Novaheim") → logit biases for "Novaheim"
|
||||
3. Store: (key=h, value=logit_biases) in memory bank
|
||||
```
|
||||
|
||||
### Recall (similarity search + logit injection)
|
||||
|
||||
```
|
||||
Query: "The capital of Zyphraxia is"
|
||||
|
||||
1. backbone(query) → h_q (896-dim vector)
|
||||
2. cosine_sim(h_q, stored_key) = 1.000
|
||||
3. Inject: logits += stored_logit_biases (per-position)
|
||||
4. Generate: "Novaheim, a city of 100..."
|
||||
```
|
||||
|
||||
### Why hidden states, not text (like RAG)
|
||||
|
||||
RAG stores text, re-encodes it each time, consumes context window.
|
||||
We store the backbone's own internal representation — no re-encoding, no context consumption.
|
||||
|
||||
## Files
|
||||
|
||||
```
|
||||
python/epimem.py ← standalone reproduction (~200 lines)
|
||||
export_onnx.py ← ONNX export from HuggingFace
|
||||
paper.md ← full paper
|
||||
results/memory_bank.json ← example: 896-dim hidden-state vectors
|
||||
schema/isis.fbs ← FlatBuffer schema for memory bank
|
||||
schema/organism.fbs ← FlatBuffer schema for organism state
|
||||
models/tokenizer/ ← Qwen 2.5 tokenizer files
|
||||
```
|
||||
|
||||
## Citation
|
||||
|
||||
```bibtex
|
||||
@article{niemi2026clivewearing,
|
||||
title={Solving the Clive Wearing Problem: One-Shot Episodic Memory for Frozen Transformers},
|
||||
author={Tommi Niemi},
|
||||
year={2026},
|
||||
organization={Rotko Networks},
|
||||
url={https://git.rotko.net/tommi/epimem},
|
||||
}
|
||||
```
|
||||
|
||||
## License
|
||||
|
||||
MIT
|
||||
Reference in New Issue
Block a user