Two ways to remember the same stream of symbols. One keeps a squeezed summary that never grows. The other keeps a record of every symbol and looks things up. Feed them both and watch where each one wins.
feed a symbol — both memories take the same step
The Pond
fixed-size state · h ← Λh + Be
— held-out
teach it first
memory held: 48 numbers — never grows
The Notebook
one entry per symbol · look up the useful one
— held-out
teach it first
memory held: 0 numbers — grows with every symbol
what the notebook learned: how far back to look
After teaching, this curve spikes at exactly the distance you asked for. That spike is the lookup rule — "the answer is K entries back." The pond has no equivalent: it cannot point at a position, only carry a blend forward.
What you should find
Set K small — say 4 — and teach both. They tie at 100%. The past is recent, so a squeezed summary still holds it.
Now push K to 20, then 30, and teach again. The notebook stays at 100%. The pond slides. Not because it is badly built, but because a fixed number of numbers cannot keep an unlimited number of separate details. As the answer moves further back, more symbols have been stirred in on top of it.
K steps back
Pond (24 dials)
Notebook
4
~98%
100%
10
~87%
100%
20
~59%
100%
30
~50%
100%
Measured by this page, at its own teaching budget; your run will vary by a few points. Chance is 25%. Given longer teaching the pond improves somewhat — but the shape holds: it falls away with distance while the notebook does not.
So why not always use the notebook?
Look at the two memory counters while you stream. The pond's stays put forever. The notebook's climbs with every symbol — and so does the work of searching it, because each new symbol must be compared against every entry so far. Double the conversation and you roughly quadruple that work.
That is the whole trade, and neither side wins outright:
Pond
Notebook
memory
fixed
grows with length
work per symbol
constant
grows with length
exact recall
fades with distance
reliable
long streams
cheap forever
eventually unaffordable
Which is why the strongest current designs are hybrids: mostly pond layers for cheap continuity, plus an occasional notebook layer for the moments when you need an exact lookup. You get the flat memory cost almost everywhere and buy back precise recall where it matters.
Honest limits of this page
The task is easy. Four symbols, one fixed distance, chance 25%. Real memory problems ask many different questions about many different details, with distances the model could not prepare for. That is where the fixed-state limit bites much harder than it does here.
Both sides are one layer. A real model stacks dozens with nonlinear blocks between them. The mechanisms shown are faithful; the scale is not.
The notebook here is a single attention head with a learned distance preference. Real attention uses many heads and content-based matching, which is more capable and more expensive than what you see.
The limitation this page demonstrates — Jelassi, Brandfonbrener, Kakade & Malach, "Repeat After Me: Transformers are Better than State Space Models at Copying," ICML 2024. arXiv:2402.01032. Proves a fixed-size state is fundamentally limited at copying and retrieval while a two-layer transformer is not.
The pond side — Gu, Goel & Ré, "Efficiently Modeling Long Sequences with Structured State Spaces" (S4), ICLR 2022. arXiv:2111.00396; and Gu & Dao, "Mamba," 2023. arXiv:2312.00752. Stability parameterization from S4D, arXiv:2206.11893.
The notebook side — Vaswani et al., "Attention Is All You Need," NeurIPS 2017. arXiv:1706.03762. The learned distance preference used here is a relative-position bias, as in Raffel et al. (T5), arXiv:1910.10683.
Not from the literature. The pond and notebook metaphors are an exposition choice. The numbers in the tables come from this page's own code on its own deliberately easy task — they illustrate a real, published effect, but they are not a benchmark and should not be quoted as one. By Daniella M. LaGuerre, August 2026.