State and checkpoints in LangGraph
How a LangGraph agent remembers — reducers decide how updates combine, and checkpoints save progress so an agent can pause, resume and survive a crash.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
State is the bag of information a LangGraph agent carries between steps, and checkpoints are saved copies of that bag, so the agent can stop and continue later.
Think of a bank passbook. Every time you visit, the printer adds new lines under the old ones — the history grows, nothing is erased. But the balance box does not grow. Each printing replaces the old balance with the new one.
That one passbook holds both behaviours: entries that pile up, and a value that gets overwritten. A LangGraph agent's state works the same way, and you choose the behaviour for each field.
Why it exists
The previous lesson built agents that finish in one go. Real agents cannot live like that.
A support conversation happens over many messages, with gaps. A research agent might run for ten minutes and crash at minute nine. A refund agent should stop and wait — possibly for hours — while a human approves the payment.
All three problems have the same shape: the agent's progress must survive a pause. Saving the bag after every step solves it. Those saved copies are called checkpoints.
How it works
step 1 ──► step 2 ──► step 3 ──► ...
│ │ │
▼ ▼ ▼
[saved] [saved] [saved] ← one checkpoint after every step
│
power cut? crash? waiting for a human?
│
▼
reload the last saved bag, continue from step 3One more idea completes the picture. Each conversation gets its own thread — its own private chain of checkpoints, found by an id you choose. Ravi's chat and Meera's chat are different threads. Same agent, separate passbooks.
Where you have already seen this
- WhatsApp shows your full chat history when you reopen it. The history is the thread.
- A half-filled government form online: you log back in and your entries are still there.
- IRCTC holds your seat selection while you go and complete the payment step.
The honest part
Deciding what belongs in the bag is design work, and it is easy to get wrong. Too little state and the agent forgets things it needs. Too much and every step drags a suitcase of old data through the graph. Most working agents keep a short list: the conversation so far, the current goal, and a few results. When in doubt, start small.
Remember this
- State travels through the graph; each field either piles up or gets replaced — you choose.
- A checkpoint is the state saved after every step; a crash or pause costs you nothing.
- A thread id keeps different users' conversations in different passbooks.
What to learn next
- CrewAI — role-based agent teams — a different framing: agents as colleagues with jobs.
- MCP — the Model Context Protocol — standard tool plumbing for any agent.
- Context window — why unbounded memory eventually breaks the model call itself.
Developer — Code and libraries.
Setup
pip install langgraphWritten against langgraph 1.2. Everything below runs offline — no model, no API key, same output on your machine.
Reducers: how updates combine
By default, a returned value replaces the old one. Attach a reducer — a function that combines old and new — and the field accumulates instead.
from typing import Annotated, TypedDict
from operator import add
from langgraph.graph import StateGraph, START, END
class State(TypedDict):
log: Annotated[list[str], add] # add = old list + new list: appends
step: int # no reducer: the new value replaces the old
def fetch(state: State) -> dict:
return {"log": ["fetched the order"], "step": 1}
def check(state: State) -> dict:
return {"log": ["checked the stock"], "step": 2}
builder = StateGraph(State)
builder.add_node("fetch", fetch)
builder.add_node("check", check)
builder.add_edge(START, "fetch")
builder.add_edge("fetch", "check")
builder.add_edge("check", END)
result = builder.compile().invoke({"log": ["started"], "step": 0})
print("log: ", result["log"])
print("step:", result["step"])log: ['started', 'fetched the order', 'checked the stock'] step: 2
Annotated[list[str], add] reads as: this field is a list of strings, and updates are combined with operator.add — old list plus new list. Both nodes returned a one-item list; the log kept all three entries. step had no reducer, so 2 overwrote 1, which overwrote 0.
For chat agents, LangGraph ships a ready-made reducer called add_messages (and a ready-made MessagesState that uses it). It appends new messages and updates a message in place when the id matches — the standard choice for conversation history.
Checkpoints and threads
Pass a checkpointer to compile() and the graph saves state automatically. Pass a thread id to invoke() and the graph knows which saved conversation to continue.
from typing import Annotated, TypedDict
from operator import add
from langgraph.graph import StateGraph, START, END
from langgraph.checkpoint.memory import InMemorySaver
class State(TypedDict):
messages: Annotated[list[str], add]
reply: str
def respond(state: State) -> dict:
n = len(state["messages"])
return {"reply": f"That is message number {n} from you."}
builder = StateGraph(State)
builder.add_node("respond", respond)
builder.add_edge(START, "respond")
builder.add_edge("respond", END)
# the checkpointer saves state after every step, keyed by thread_id
graph = builder.compile(checkpointer=InMemorySaver())
ravi = {"configurable": {"thread_id": "ravi"}}
meera = {"configurable": {"thread_id": "meera"}}
print(graph.invoke({"messages": ["hello"]}, ravi)["reply"])
print(graph.invoke({"messages": ["how are you"]}, ravi)["reply"])
print(graph.invoke({"messages": ["namaste"]}, meera)["reply"])
saved = graph.get_state(ravi)
print("ravi's saved messages:", saved.values["messages"])That is message number 1 from you. That is message number 2 from you. That is message number 1 from you. ravi's saved messages: ['hello', 'how are you']
Read that output again. The second call passed only "how are you", yet the agent knew it was message number two. The checkpointer loaded Ravi's saved state, the add reducer appended the new message, and the node saw both. Meera's thread started fresh at one. Nothing about memory appears in the node code — persistence is configuration, not logic.
InMemorySaver lives in RAM and dies with the process. For real use, the interface is identical with a database behind it — pip install langgraph-checkpoint-sqlite gives you SqliteSaver, and there is a Postgres version. Swap the saver, change nothing else.
Pausing for a human
Because state survives between calls, a graph can stop mid-run and wait — the foundation of approval steps.
from typing import TypedDict
from langgraph.graph import StateGraph, START, END
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.types import interrupt, Command
class State(TypedDict):
amount: int
status: str
def approve(state: State) -> dict:
# pause the run and ask a human; the answer comes back as the return value
answer = interrupt(f"Refund Rs {state['amount']} - approve? (yes/no)")
return {"status": "paid" if answer == "yes" else "rejected"}
builder = StateGraph(State)
builder.add_node("approve", approve)
builder.add_edge(START, "approve")
builder.add_edge("approve", END)
graph = builder.compile(checkpointer=InMemorySaver())
cfg = {"configurable": {"thread_id": "refund-77"}}
paused = graph.invoke({"amount": 500}, cfg)
print("paused with question:", paused["__interrupt__"][0].value)
finished = graph.invoke(Command(resume="yes"), cfg)
print("after human answered:", finished)paused with question: Refund Rs 500 - approve? (yes/no)
after human answered: {'amount': 500, 'status': 'paid'}The first invoke stopped at interrupt() and returned the question under __interrupt__. The process could exit here, and the thread would wait in the checkpointer for days. The second invoke sends Command(resume="yes") into the same thread, and the paused interrupt() call returns that value as if nothing had happened. Interrupts need a checkpointer — there is nowhere to park the paused state without one.
Common mistakes
Using a checkpointer without a thread id. The graph has no idea which passbook you mean:
ValueError: Checkpointer requires one or more of the following 'configurable' keys: thread_id, checkpoint_ns, checkpoint_id
The fix is the config argument: {"configurable": {"thread_id": "..."}} on every invoke.
One thread id for everybody. Give all users "main" and every user sees everyone else's history merged together. The thread id must come from your application — the user id, the ticket number, the chat session.
Expecting InMemorySaver to survive a restart. It is a Python dictionary. Restart the process and every thread is gone. It exists for development and tests; production wants the SQLite or Postgres saver.
Appending forever. An add reducer never forgets, so a long conversation grows without limit — and the whole state is serialised at every checkpoint, every step. Cap it (keep the last N messages, summarise the rest) before it becomes a context window problem and a storage bill at the same time.
Try it yourself
After the three invoke calls in memory_demo.py, add:
for snap in graph.get_state_history(ravi):
print(snap.values)You will see every checkpoint of Ravi's thread, newest first, back to the empty start. This is time travel — you can re-run the graph from any of those snapshots, which is how you debug an agent that went wrong at step 3 of 20 without paying for steps 1 and 2 again.
What to learn next
- CrewAI — role-based agent teams — a different framing: agents as colleagues with jobs.
- MCP — the Model Context Protocol — standard tool plumbing for any agent.
- Context window — why unbounded memory eventually breaks the model call itself.
Researcher — Mathematics and papers.
Checkpointing as a write-ahead log
A checkpointer implements a versioned key-value log with the schema: (thread_id, checkpoint_ns, checkpoint_id) → serialized channel values + versions + pending writes. Where:
thread_id— the user-chosen partition key; one thread is one totally-ordered history.checkpoint_ns— a namespace distinguishing subgraph checkpoints from their parent's.checkpoint_id— a monotonically sortable unique id per superstep, giving the total order.
One checkpoint is written per superstep, after all updates for that step are reduced. The recovery guarantee is the database one: on crash, replay resumes from the last committed checkpoint. Node side effects, however, are not transactional with the checkpoint. A node that fired a payment API and crashed before its superstep committed will fire it again on resume. Exactly-once state, at-least-once side effects — the same contract as most stream processors, and the reason payment-like calls belong behind idempotency keys even inside a checkpointed graph.
Cost model: per superstep, one serialisation of changed channels plus one write; per resume, one read and deserialisation of the latest checkpoint. State size directly multiplies both, which is the formal version of "do not carry a suitcase".
Reducers as merge operators
Requiring each channel's updates to merge through a binary operator $r(v_{old}, v_{new})$ is what makes concurrent branches well-defined. If parallel nodes $A$ and $B$ both write channel $c$ in one superstep, the result is $r(r(v, u_A), u_B)$ up to ordering — deterministic for associative, commutative $r$ (list concatenation is associative but not commutative; LangGraph fixes update order by node ordering to keep runs reproducible). A channel without a reducer that receives two writes in one superstep raises an invalid-update error instead of guessing. This is the dataflow lineage of the design: channels are single-assignment-per-step variables, reducers are the join.
add_messages is a domain-specific reducer: append by default, replace on matching message id, enabling in-place message edits and deletions through sentinel values — an operational-transform flavour rather than a pure monoid.
Time travel, forking and replay
get_state_history exposes the full checkpoint chain; invoking with a checkpoint_id in the config replays from that snapshot, and writing a modified state with update_state before resuming forks the history. Determinism on replay holds only if node functions are pure given state — any node reading the wall clock, a random seed, or a live API can diverge. The debugging methodology this enables mirrors record-replay debuggers rather than logging: reproduce, rewind, poke, re-run.
Durable execution in context
The pause-resume mechanism places LangGraph in the durable execution family — Temporal, Azure Durable Functions, AWS Step Functions — where a workflow's progress is externalised so the process is disposable. The distinguishing constraint in the agent setting is that control flow is chosen by a model at run time, so the "workflow definition" cannot be statically validated the way a Step Functions state machine can. Human-in-the-loop interrupts are the same mechanism as crash recovery, deliberately: an approval is a pause with a payload, not a special case.
For the human-oversight patterns built on interrupts — approval gates, edit-before-execute, escalation — the literature keyword is human-in-the-loop agent supervision; a dedicated lesson in this section covers the design side.
Sources
- Malewicz et al., 2010, Pregel: A System for Large-Scale Graph Processing — the superstep model checkpointing slots into.
- Valiant, 1990, A Bridging Model for Parallel Computation — bulk synchronous parallelism.
- Shapiro et al., 2011, Conflict-free Replicated Data Types — the merge-operator reasoning behind reducer design.
- LangGraph persistence documentation — the checkpointer interface is defined by the framework, not a paper; verify details against your installed version.
What to learn next
- CrewAI — role-based agent teams — a different framing: agents as colleagues with jobs.
- MCP — the Model Context Protocol — standard tool plumbing for any agent.
- Context window — why unbounded memory eventually breaks the model call itself.