← Back to all posts

The Brain Is a Key, Not a Box

The storage metaphor of memory — encode, compress, file, retrieve, unzip — fails for brains and for neural networks in the same three ways: nothing is stable, nothing is located, and nothing survives being switched off. Victoria Trumbull's argument that the brain enables memory rather than containing it is a description of parametric memory in machine learning; her two readings of 'time stores memory' map onto recurrence and attention; and CatBoost's ordered target statistics — which the paper calls an 'artificial time' — are the clearest small example of a model that keeps the past in an ordering instead of compressing it into a value.

memoryphilosophy-of-mindneural-networksparametric-memoryengramreconsolidationbergsondurationeternalismenactivismcatboostxgboostgradient-boostingordered-boostingragcontext-windowhallucinationessay

There is a metaphor we all carry around and almost never examine. It goes like this:

experience ──▶ encode ──▶ compress ──▶ store in a folder ──▶ retrieve ──▶ unzip
             (hippocampus)            (a "memory trace")        (recall)

You experience something, your brain encodes it, compresses it like a zip file, files it in a folder — hippocampus, cortex, somewhere — and later you open it and watch it play. Nothing about that feels like a philosophical position. It feels like the obvious description of remembering.

It is a very specific philosophical position with a name: the storage metaphor of memory, encoding a kind of reductive materialism in which the memory is a physical thing located inside the brain — what neuroscience calls an engram. The metaphor is older than neuroscience (Plato's wax block, Aristotle's seal in wax, then traces, vibrations, switches), and for most of that history it was not a prediction about biology so much as the only vocabulary anyone had (Danziger; SEP: Memory).

The philosopher Victoria Trumbull argues it fails, and — more interestingly — that it has misdescribed the shape of the answer: the brain does not contain memories, it enables them, the way a radio enables music without containing the music (Memories are not stored in the brain).

Victoria Trumbull — "Memories are not stored in the brain" (IAI, 2026)

What follows takes that seriously without either swallowing it or dismissing it — and then asks the question this blog always ends up asking, because we built machines out of the same metaphor and it is breaking in the same three places.

Three ways the zip file metaphor breaks

1. Recollection rewrites the file. A zip file does not change when you unzip it. A memory does. The mechanism is called reconsolidation, and it was pinned down in a famous experiment: a consolidated fear memory in rats, once retrieved, becomes temporarily unstable and needs new protein synthesis to restabilize; block the synthesis and it does not come back intact (Nader, Schafe & LeDoux, Nature, 2000). Retrieval is not read-only. It is a rewrite followed by a save.

That is the physiological version of what psychology had documented for decades: Bartlett's subjects reshaped stories toward their own expectations (1932); Loftus and Palmer's subjects "remembered" broken glass that was never there after being asked about cars that smashed into each other (1974); and constructive-memory research now treats remembering and imagining as overlapping operations on the same machinery (Schacter & Addis). A storage system whose retrieval corrupts the store is not a filing cabinet. It is closer to a wiki that rewrites itself every time someone reads a page.

2. There is no folder. Sixty years of looking for the location of a specific memory did not produce a box. Engrams are real — cells have been identified, tagged, artificially activated to trigger a behavior, and silenced to prevent one (Josselyn & Tonegawa, Science, 2020) — but the trace is not addressable the way the metaphor needs it to be. One memory is represented by ensembles distributed across multiple brain regions, each contributing a piece, with the same cells participating in overlapping memories (Roy et al., 2022). Karl Lashley spent decades cutting into rat brains looking for the trace, found that removing more tissue degraded performance smoothly instead of deleting memories one at a time, and titled the write-up In search of the engram (reviewed in Josselyn, Köhler & Frankland, 2015).

3. The store is never powered off. A hard drive can sit on a shelf for a decade and return the same bits. A brain spends roughly 20% of the body's energy budget (Raichle & Gusnard, 2002), most of it on activity unrelated to the task in front of it — the spontaneously active, task-negative network called the default mode (Raichle et al., 2001). Nothing is parked. When activity stops, the memory does not wait on the shelf; the system that could reconstitute it is gone.

So we are describing an object with none of the properties of a file: it changes on read, has no location, and exists only while running. A radio does not contain Beethoven's Seventh; it is a structure that, powered and tuned, produces it. Smash the radio and you have not destroyed the symphony — you have destroyed one way of hearing it.

What "time itself stores memory" could mean

The radical half of the argument is that memory is not stored anywhere, because the thing doing the storing is time. Two traditions get compressed into that sentence, and they are worth separating.

Reading A: Bergsonian duration. In Matter and Memory (1896), the brain is not an archive but "an organ of attention to life" — a filter that narrows the whole of the past down to what is useful now (Bergson; SEP: Bergson). The past is not gone, it is simply not useful at the moment. Memory is duration: a living accumulation that persists rather than a shelf that survives. Remembering is less like retrieving a document than like turning your attention to a region of your own temporal extent.

Reading B: the block universe. Under eternalism every moment of spacetime exists equally; 1900 and 2026 are both real at different coordinates, the way New York and Paris are both real at different places. That is one live option among several in the philosophy of time (SEP: Time; Sider's four-dimensionalism), with presentists denying the whole picture. On this reading a memory needs no storage location: the event is still at its coordinate, and the brain's job is to be a structure complex enough to re-establish a relation to it.

Both readings land on the same inversion: the brain is an enabling condition, not a storage location — and we mistook one for the other because the brain is the thing we can watch working while remembering happens.

Where the argument is harder than it looks

Two steps are weaker than the rhetoric suggests, and both matter for what follows.

The physics does not deliver the conclusion by itself. Eternalism is an interpretation, not a result; relativity removes a privileged global "now" while the debate between presentism, the growing block, and the moving spotlight continues. Rovelli's The Order of Time goes further and then lands somewhere that cuts against the simple "the past is just elsewhere" picture: the arrow of time, he argues, is something we recover from the records the world leaves behind (Rovelli). If traces give time its arrow, then "time stores memory" quietly becomes "memory is what makes time storable" — stranger, and more interesting.

Even granting the block universe, you still need a key. If the event at t₁ exists, what makes a brain at t₂ about that event rather than any other coordinate? Something has to connect them — a causal chain, a structural correspondence, a disposition to respond to the right thing. Whatever that is, it is a trace, and traces are what the causal theory of memory is built on (SEP: Memory). So the storage metaphor is better called misplaced than false: the trace is not the repository, it is the hook that lets a present process re-instantiate a past one. That is the structure the extended and enactive traditions build on (Clark & Chalmers; Varela, Thompson & Rosch; SEP: Embodied Cognition).

The empirical bridge is worth one line: divers who learned word lists underwater recalled them better underwater (Godden & Baddeley, 1975). The room is part of the key, which the box model cannot represent. And a caution: Whitehead's Process and Reality is a metaphysics in which the past is causally efficacious in the present, not a claim about neurobiology. Reading it as physics is how a good idea becomes a bad one.

The same metaphor, in the machine

The storage metaphor was also the founding assumption of AI, and here the philosophy stops being a spectator sport.

Classical symbolic AI was a database. Knowledge sat in rows — Paris = capital_of(France) — in semantic networks and ontologies, and reasoning was retrieval plus rules. Connectionism replaced that with something stranger. A neural network has no rows; it has a pattern of connection strengths, and a "fact" exists only as a disposition distributed across millions of weights, re-instantiating when the right input arrives. There is no folder in GPT for "Paris" — no address you can point to and say here it lives.

This is standard vocabulary rather than a poetic reading: we call it parametric memory, knowledge held in weights, as opposed to non-parametric memory, knowledge held where you can look it up (Petroni et al., 2019). Which means Trumbull's sentence — the system enables memory without containing it — is a description of a trained network, not an analogy for one. The AI form of reductive materialism says intelligence is in the weights; the everyday experience of a neural network is that nothing is in anything.

The evidence that there is no folder is that we keep failing to edit one. If a fact were a row, changing it would be an UPDATE statement; instead, locating and editing a single factual association means finding the mid-layer activations that carry it, and the edit remains entangled with everything else those weights do (Meng et al., ROME, 2022). Distributed is not a slogan here; it is a budget line.

The failure mode matches too. A generative model does not retrieve a stored sentence — it reconstructs a continuation, and when the reconstruction has no grounding the result is fluent and wrong, structurally rather than accidentally (Xu & Jain; for the human taxonomy, Ji et al.). Confabulation is what reconstruction looks like when the key turns in the wrong lock.

The analogy should not be oversold: human confabulation and LLM hallucination share a shape — construction without verification — but not a mechanism, one running through reconsolidation in an organism with stakes and the other through next-token distributions over a corpus. The useful claim is structural. Both are keys, and neither is a box.

The two readings, in code

Machine learning has spent thirty years building one architecture for each of Trumbull's two readings, and arguing about which is right.

Recurrence: memory lives in the process, not the weights. A recurrent network carries a hidden state forward — the past exists as how it has changed the present state:

h_t = f(h_{t-1}, x_t)   # the past is not retrieved; it persists

That is duration in code: nothing is looked up, the past is present as accumulated modification. The disanalogy is the second half of the story. A hidden state is a fixed-width compression of the past, not a retention of it; Bergson's duration keeps everything and prioritizes by attention, while a vanilla RNN forgets and its gradient signal decays over long horizons (Bengio et al., 1994). Gates, and now selective state-space models like Mamba, are negotiations with exactly that limit.

Attention: leave the past where it is and learn to point at it. Transformers hold the whole sequence and let every position look at every other, recomputing each token's representation from the entire context on every forward pass. Nothing is zipped, because nothing is unzipped. Two corrections to the popular telling: attention is content-addressed rather than coordinate-addressed (positions enter as added information, not as the mechanism), and the KV cache is an artifact of a single generation, recomputed next time and present only while the process runs. The past there is held, not kept.

The engineering trajectory is the interesting part. We spent years trying to store everything inside the model; the last several years have gone the other way, and the results keep favouring it — RAG, nearest-neighbour LMs (whose title, Generalization through Memorization, is the irony of the decade), RETRO at trillion-token scale, Memorizing Transformers, and now the entire industry of vector databases and long context. That is Trumbull's move, executed by engineers who have never read Bergson: do not force the model to contain the past — leave it out in the world and make the model a structure that can establish a relation to it.

With the same caveat a brain has: having the past laid out is not the same as using it. Retrieval quality degrades sharply depending on where the relevant passage sits in a long window, worst in the middle (Liu et al., 2023). The key has to fit the lock. Storage was never the hard part.

Case study: the zip file in gradient boosting

The cleanest small example is gradient-boosted trees, and specifically the difference between XGBoost and CatBoost — usually explained as a feature-engineering advantage, actually a disagreement about where memory should live.

Start with the naive move, which is the zip file as a line of code. For a categorical column like city, the shortcut is mean encoding:

# the zip-file move: compress all of history into one stored number
df["city_encoded"] = df.groupby("city")["target"].transform("mean")

Every Paris row now carries one number computed from every Paris row — including rows that come after it, and including its own label. The past has been compressed into a constant stored in a cell, and the future has leaked into it. That is the storage metaphor as a bug: a value that is the memory, true regardless of when you read it.

CatBoost's answer is the philosophical one, and the paper says it almost openly. Instead of computing the encoding once and storing it, it computes, for each example, a statistic from the examples that come before it in an ordering — and to make that possible in a static dataset, the authors introduce what they call an artificial "time":

"Clearly, the values of TS for each example rely only on the observed history. To adapt this idea to standard offline setting, we introduce an artificial 'time', i.e., a random permutation σ of the training examples." — Prokhorenkova et al., CatBoost, 2018

To avoid storing a memory as a value, give the data a time — and the encoding becomes a relation to a history instead of a constant. Extended to gradients, the same mechanism is ordered boosting, which prevents prediction shift: the leakage you get when the same rows both fit the model and estimate what it should predict.

Now the correction that keeps the analogy honest, because it is tempting to read "ordered" as "temporal" and stop. The paper's own word is artificial: by default that ordering is a random permutation, several of them across boosting steps. Your event stream's chronology is not what the library is respecting, which is why CatBoost exposed a has_time switch documented as: "Use the order of objects in the input data (do not perform a random permutation of the dataset at the preprocessing stage)… Default: FALSE (not used; permute input dataset)" (R package source).

So ordered means ordered for anti-leakage purposes, not chronological. Those are different guarantees, and conflating them is the mistake I warned about last post: label leakage is prevented structurally by the library, while temporal leakage is your responsibility, prevented by time-ordered folds and point-in-time-correct features. A random permutation is a beautiful trick against one and a quiet hazard for the other.

With that caveat, the reading holds. CatBoost also builds oblivious trees — the same split at every level, which makes them fast to evaluate and, in the paper's phrase, "less prone to overfitting." You cannot open a trained model and find row #4521 in it: the training set's influence is spread across thousands of splits and leaf values, enabling prediction without containing data. XGBoost deserves no strawman here — it is equally distributed, and has handled categoricals natively since 1.5 with partition-based splits (docs). Historically it asked you to bring your own encoding, which meant carrying the leakage risk yourself. That is a difference in where the history lives during training, not between a database and a brain.

The honest version: both are keys. What differs is whether the history is precomputed into a column and treated as timeless — the zip file — or recomputed per example from an ordering that defines what counts as the past. CatBoost's contribution to this essay is that it is ordinary, popular software whose designers independently decided the past should live in an order rather than a value.

What this means for how we build

A warning first: machines really do store. A file on disk is exactly the box model, and for silicon the storage metaphor is accurate rather than metaphorical. The takeaway is narrower — do not assume a brain is doing what your database does, and be suspicious when a cognitive claim arrives pre-shaped like a system you already know how to build.

The suspicion runs both ways, because "memory" in agent harnesses is almost always a storage layer: embeddings as engrams, a vector database as the hippocampus, retrieval as unzip-and-read. If the philosophy is right, that imports the failure modes:

  • Retrieved memories get treated as facts. A file is authoritative; a reconstruction is a hypothesis. If recall is construction, then a system that learns from its own recalled state is compounding its own edits — the dynamic behind drift and confidently retrieved falsehood.
  • The room gets dropped. Context-dependence means the situation is part of the memory; retrieval keyed on text similarity alone discards state, task, provenance, and what the system was doing when it learned the thing.
  • Architecture assumes a powered-off shelf. A brain never has one. A log you can re-derive from is closer to that than a store you load from — the event log is the honest memory, narratives are reconstructions built from it, a durable daemon survives the power cycle.

The economics follow. If the storage metaphor is wrong, then its plan — make the container bigger so it can hold more — is a plan for buying hard drives. The evidence supports a smaller structure with better access: less memorization, more reconstruction skill, memory as a relation between system and world rather than content inside the system. That is the same thesis this blog keeps arguing from the model side, where specialization, not size, is the lever and the deployable unit is the model plus the harness.

And the honest limit: none of this is a design specification. Philosophy can tell you the storage metaphor is a bad map of biological memory and a leaky one for models. It cannot tell you how to build better agent memory, and "the past still exists at its coordinate" is not a caching strategy. What it can do is change your default assumption and the failure mode you watch for — memories confidently retrieved, unaccountably shaped by context, and quietly rewritten every time you read them.

What this changed in my view

Three things moved while I was working through this, and one didn't.

I had assumed the engram debate was settled empirically, in the storage metaphor's favour. Find the trace, tag it, activate it — Tonegawa-style experiments really do this, and I had filed it as "the box exists, the philosophers are quibbling about vocabulary." The distributed-engram work changed that: the trace exists and is not a location. That is a result the storage metaphor cannot represent, which is stronger than a philosophical preference.

I had treated "memory is reconstruction" as a slogan about false memory. Reconsolidation makes it a mechanism — recall destabilizes, edits, re-saves — and a mechanism is a warning about any system that reads its own memory and writes back, including the agents we are building.

The best example turned out to be small and ordinary. I expected a large language model, where the story is diffuse and hard to verify. Instead it was seventy lines of gradient boosting: CatBoost's paper introduces an artificial time to avoid storing a memory as a value, and hands you a switch to say "no, use the real order." I nearly wrote the tidy version of that — that CatBoost respects time. It does not; it respects an order, usually a random one, and ordered is about leakage rather than chronology. The correction improved the essay, because it separates two things the storage metaphor also conflates: keeping the past around and keeping it in order are different achievements, and only one of them is free.

What did not move: brains are physical, engrams are real, weights are real, and the storage metaphor is not worthless. It is just load-bearing in the wrong place — which is the most consequential kind of wrong a metaphor can be.

References

The primary argument

  • Trumbull, V. Memories are not stored in the brain — Institute of Art and Ideas talk, 2026. The claim under discussion: the brain enables rather than contains, and the storage metaphor mistakes an enabling condition for a storage location. Figure above links to the talk.
  • Michaelian, K. and Sutton, J. Memory — Stanford Encyclopedia of Philosophy. Traces, the causal theory of memory, distributed and extended accounts, and the objections to each.
  • Semon, R. The Mneme (1921; German original 1904) — the source of the term engram, and of ecphory: Semon already distinguished the trace from the act of reactivating it, which is a reminder that the box was never the whole theory.
  • Danziger, K. Marking the Mind: A History of Memory — Cambridge University Press, 2008. How the storage vocabulary — wax, seals, traces, vibrations, switches, files — was carried across four centuries of theories.

The three failures

Time, duration, and the block universe

Parametric and non-parametric memory in machines

Where the past lives in gradient boosting

Minds that are not sealed containers

This blog, on memory and logs