Agent memory
How agents keep hold of facts after the conversation scrolls away — short-term history, rolling summaries, and long-term notes the agent saves on purpose.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Agent memory is everything an agent can still use after the words have scrolled out of the conversation.
Think of a kirana shopkeeper who gives regular customers credit. He chats with you at the counter, and that chat lives in his head. But your running balance goes into the khata — the notebook — because his head cannot hold every customer's account.
Next month he does not try to remember. He opens the notebook.
Why it exists
A language model has no memory between conversations. None. Each new chat starts from a blank page.
Within one conversation, it has only the context window — the fixed amount of recent text it can see, explained in context window. When a chat grows past that limit, the oldest words fall off the edge.
For a chatbot answering one question, fine. For an agent — a model doing long, multi-step work, see AI agents — this is crippling. An agent that forgets your city, your order number, or its own earlier decisions cannot finish a real job.
So agents get memory the same way the shopkeeper does. Not a bigger head — a notebook.
The three layers
┌── recent turns (kept word for word)
prompt ←──┼── rolling summary (older turns, compressed)
└── saved notes (facts written down on purpose,
searched when needed)- Short-term: the last few turns, kept exactly. This is the counter conversation.
- Summary: when turns get old, they are squeezed into a short recap so the gist survives.
- Long-term notes: facts the agent deliberately writes down — your city, your preferences, a decision made last week. Stored outside the model, found again by search. The searching works like RAG: look it up first, then answer.
Where you have already seen it
ChatGPT-style assistants sometimes say "I'll remember that" — a note being written. Your food delivery app greets you with your saved address instead of asking again. And the shopkeeper's khata has worked this way for centuries.
The honest part
The genuinely unsolved problem is deciding what to save. Save too little and the agent forgets your city. Save too much and the notebook becomes a haystack where search returns the wrong needle. Every serious agent team is still hand-tuning this, and no framework has made the problem disappear.
Remember this
- Models forget everything between chats; memory lives outside the model.
- Recent turns stay exact, old turns become a summary, key facts become notes.
- The hard part is choosing what to write down, not where to store it.
What to learn next
- Context window — the limit that forces all of this.
- What is RAG? — the retrieval machinery agents point at their own notes.
- Human-in-the-loop agents — pausing safely requires remembering the plan.
Developer — Code and libraries.
Setup
python3 --version # standard library only — nothing to installFramework-free on purpose. Every agent framework's memory feature is a variation of the fifty lines below, and seeing the moving parts once makes their documentation readable.
The three layers in runnable form
BUDGET = 45 # word limit for raw history kept in the prompt
history = [] # short-term: the turns of this conversation
summary = [] # rolling gist of turns that fell out of the window
notes = {} # long-term: facts saved on purpose, survives restarts
def add_turn(speaker, text):
history.append(f"{speaker}: {text}")
while sum(len(t.split()) for t in history) > BUDGET:
oldest = history.pop(0)
# A real agent asks the model for a one-line summary here.
summary.append(" ".join(oldest.split()[:5]) + "...")
def build_prompt(question):
parts = ["EARLIER (compressed): " + " | ".join(summary)] if summary else []
parts += history
facts = [f"{k} = {v}" for k, v in notes.items() if k in question.lower()]
if facts:
parts.append("SAVED NOTES: " + "; ".join(facts))
parts.append(f"user: {question}")
return "\n".join(parts)
add_turn("user", "Hi, I want a mechanical keyboard delivered to my hostel in Pune")
notes["city"] = "Pune" # the agent chose to save this fact the moment it appeared
add_turn("agent", "Nice choice. Which switches do you prefer, clicky or silent ones?")
add_turn("user", "Silent ones please, my roommate studies late at night most days")
add_turn("agent", "Understood, silent switches it is. Anything else I should know today?")
prompt = build_prompt("What will delivery cost to my city?")
print(prompt)
print("---")
print("words in prompt:", len(prompt.split()))EARLIER (compressed): user: Hi, I want a... agent: Nice choice. Which switches do you prefer, clicky or silent ones? user: Silent ones please, my roommate studies late at night most days agent: Understood, silent switches it is. Anything else I should know today? SAVED NOTES: city = Pune user: What will delivery cost to my city? --- words in prompt: 56
Read that output — it contains the whole argument
The first turn is where the user said Pune. That turn fell over the budget and got compressed to user: Hi, I want a... — and Pune vanished from the compressed line.
Yet the final prompt still knows the city, because the agent wrote city = Pune into notes the moment it appeared. That is the entire case for deliberate memory writes in one screen: summaries lose facts, notes keep them. If the note had not been saved, no amount of cleverness at question time could recover the city.
Details worth a second look:
The while loop in add_turn trims from the oldest end until the raw history fits the budget. Real systems budget in tokens rather than words, and ask the model itself for the one-line summary — which costs a model call, so trimming has a price.
build_prompt assembles three sections in a fixed order: compressed past, exact recent turns, then relevant notes. Keeping notes close to the question helps, since models attend best to the edges of the prompt.
Note recall here is a keyword match (if k in question.lower()). Honest and readable, and it breaks the moment the user asks "what about shipping to my town?" Real systems search notes by meaning with embeddings inside a vector database — the same machinery as RAG, pointed at the agent's own notebook.
Common mistakes
Stuffing instead of managing. Appending forever until the context overflows, then wondering why the agent forgot its instructions. The system prompt itself can fall off the edge. Budget explicitly, always.
Summarising away identifiers. Order ids, amounts, emails, cities — the details summaries drop first are the details agents need most. Rule of thumb: identifiers go to notes before their turn is eligible for compression.
Trusting the model to remember across sessions. It cannot. If a fact is not in a store you control, it is gone when the process restarts. Test by killing the process and asking again.
Never updating or expiring notes. The user moves from Pune to Nagpur; the shopkeeper crosses out the old balance when you pay. Store a timestamp with every note, overwrite on conflict, and prefer the newest fact at recall time.
Try it yourself
Drop BUDGET to 25 and rerun — predict which turns compress before you look. Then add notes["roommate"] = "studies late at night" and ask a question containing "roommate". Then break it on purpose: ask "what will shipping cost to my town?" and watch the keyword match fail. That failure is the reason embeddings exist.
What to learn next
- Context window — the limit that forces all of this.
- What is RAG? — the retrieval machinery agents point at their own notes.
- Human-in-the-loop agents — pausing safely requires remembering the plan.
Researcher — Mathematics and papers.
Retrieval scoring over a memory stream
Park et al. (2023), Generative Agents (arxiv.org/abs/2304.03442), store every observation in an append-only stream and retrieve by a linear combination:
$$ \text{score}(m, q) = \alpha_{\text{rec}} \cdot \text{recency}(m) + \alpha_{\text{imp}} \cdot \text{importance}(m) + \alpha_{\text{rel}} \cdot \text{relevance}(m, q) $$
Symbols: $m$ is a stored memory, $q$ the current query context; $\text{recency}$ decays exponentially with time since last access; $\text{importance}$ is a model-assigned score of how consequential the memory is; $\text{relevance}$ is embedding cosine similarity to $q$; the $\alpha$ weights balance the three (all set to 1 in the paper). The design point: pure semantic similarity is insufficient — a memory can be relevant and stale, or recent and trivial.
Paged memory and self-editing
Packer et al. (2023), MemGPT (arxiv.org/abs/2310.08560), treat the context window as RAM and external storage as disk, with the model issuing function calls to page data in and out — memory management as tool use. The line of work continues in Letta, the successor system, as of 2026 one of several production takes alongside LangGraph's checkpointer/store split and CrewAI's memory layers. The framework APIs differ; the RAM/disk decomposition recurs everywhere.
Episodic and procedural variants
- Shinn et al. (2023), Reflexion (arxiv.org/abs/2303.11366) — after a failed attempt, the agent writes a verbal self-critique to an episodic buffer; conditioning retries on it improves success without weight updates. Memory as a gradient substitute.
- Wang et al. (2023), Voyager (arxiv.org/abs/2305.16291) — procedural memory: skills are stored as verified, reusable code indexed by embedded descriptions, and the library compounds over the agent's lifetime.
Cost analysis
Sending full history makes turn $n$ cost $O(nc)$ input tokens and a conversation of $T$ turns cost $O(T^2 c)$ in total, with $c$ the mean tokens per turn — before accounting for attention's own quadratic cost in sequence length. A rolling summary caps per-turn cost at $O(B + s)$, where $B$ is the raw-history budget and $s$ the summary length, at the price of lossy compression plus one summarisation call per eviction. Note retrieval via approximate nearest neighbour is sublinear in note count (see vector databases), which is why the notes layer scales and the stuff-everything approach does not.
Open problems
Write policy — what to store, at what granularity, with what deduplication — remains heuristic; learned write policies are an active research direction without a settled recipe. Conflict resolution between contradictory memories mostly reduces to timestamps and overwrites. Evaluation is the weakest link: long-horizon memory benchmarks are young and small relative to the claims made in framework marketing, so as of 2026, published memory features deserve the same scepticism as any unevaluated component.
What to learn next
- Context window — the limit that forces all of this.
- What is RAG? — the retrieval machinery agents point at their own notes.
- Human-in-the-loop agents — pausing safely requires remembering the plan.