Here is a fact that explains most of the frustration people have with agents. A language model holds no memory between calls. None. Every session starts from zero — same weights, blank slate, no idea who you are or what you decided last Tuesday.
When an agent seems to remember you, that is not the model remembering. That is scaffolding. Somewhere, a system wrote things down — files, logs, retrieved notes — and quietly read them back in before the model said a word. The memory was never in the intelligence. It was in the filing.
This distinction sounds academic until you notice what it does to your costs.
A while back I argued that most "stupid" agent behavior is really underbriefing — the model was never told what you know. The brief is what you hand the system at the start of a session. Memory is what survives the end of one. Different mechanism, different failure. A bad brief wastes a session. Missing memory wastes every session, because each one relearns what the last one already knew.
Watch where your corrections go. You tell the agent the tagline is never abbreviated. It complies, beautifully, until the session ends — and the correction evaporates with the context. Next week you make the same correction again, with slightly less patience. A correction made in chat is a conversation. A correction written into the system — a rules file, a decision log, a style memory the agent loads on every run — is an asset. The first one you pay for weekly. The second one you pay for once.
So the working unit of agent memory is not a bigger context window. It is a file you own. Ours are unglamorous: decisions made and why, corrections given, things we tried and killed. Every run starts by reading them. The agents on my team are not smarter than anyone else's — they are better filed. The system compounds because nothing it learns is allowed to live only in a transcript.
There is a timely reason to care. Frontier models became free downloads this month, and swapping one for another now costs roughly a lunch break. Your memory files are indifferent to that. They are model-agnostic by construction — plain files that make whatever model reads them act like it has worked for you for a year. The models keep changing. The filing is what stays yours.
Start with one file and three headings. Decisions: what you chose and why, so it is never relitigated. Corrections: every note you have given twice. Killed: what you tried, and why it died, so nothing gets reinvented. Then point your agents at it on every run. The rule that makes it stick is simple: the second time you correct the same mistake, the correction goes in the file, not the chat.
Takeaway: An agent's memory is not a feature you wait for. It is a file you write. If a lesson lives only in a chat transcript, you have scheduled yourself to teach it again. ✱
