Downloads · 30 days
0
ARotting/fast-weight-time-machine
fast-weight-time-machine is a machine learning model from ARotting. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Fast-Weight Time Machine tests temporary variable binding with an explicit, sequence-local weight matrix. A controller receives key/value writes, produces a write strength, and accumulates outer products in a fast mem…
Downloads · 30 days
0
Access
Public
Updated Jul 30, 2026
Repo size
23.4 KB
Likes
0
Public
Click a slice to open those files.
.safetensors23.4 KB · 44%
From the Hugging Face model README
Fast-Weight Time Machine tests temporary variable binding with an explicit, sequence-local weight matrix. A controller receives key/value writes, produces a write strength, and accumulates outer products in a fast memory. A later query key reads the memory in one matrix-vector operation. The learned parameters stay fixed between examples; only the fast matrix changes inside each sequence.
The historical anchor is Schmidhuber's 1992 Learning to Control Fast-Weight Memories, which described feedforward controllers producing context-dependent weight changes and included adaptive temporary-variable binding. This project is a modern outer-product interpretation tested against a similarly sized GRU. It does not claim to reproduce the original implementation or experiments exactly.
The models trained on 16,000 sequences containing four writes and 12 distractors. Each held-out condition contained 4,000 new sequences.
| Condition | Fast weights, 2,508 params | GRU, 3,050 params |
|---|---|---|
| 4 bindings, 12 distractors | 100.00% | 37.40% |
| 8 bindings, 12 distractors | 100.00% | 27.03% |
| 4 bindings, 64 distractors | 100.00% | 31.15% |
| 8 bindings, 64 distractors | 99.98% | 23.83% |
This task is intentionally aligned with the fast-weight architecture: an explicit outer-product matrix can bind a key and value, while the GRU must compress all bindings into one recurrent vector. The comparison demonstrates the inductive bias; it is not a general claim that fast weights outperform GRUs on arbitrary sequences.
uv run python projects/fast-weight-time-machine/train.py