Most agent memory demos store every message and call it a day. That works until the store fills with stale preferences, contradictory facts, and half-finished plans — and the agent starts confidently acting on them. Long-term memory is not a buffer. It is a small database your model writes to without review, and reads from without skepticism. Treat it that way and the design questions get concrete.
This is about memory that survives a thread: what a user told you last month, what a customer's account looks like, what an agent learned the last time it tried a task. Within-thread context management — compaction, summarization, what to keep in the window — is a different problem, and we covered it separately in context engineering for long-running agents.
Three kinds of memory, three different failure modes
Lumping everything into one vector store is the most common mistake. Split by what the data is, because the write and read rules differ.
Facts about the user or account. Stable, small, high-value: timezone, preferred output format, the repo they work in, their plan tier. These belong in a keyed store — one namespace per user, one row per fact — not in a similarity index. You read them by key on every run, not by embedding a query and hoping. Failure mode: contradiction. The user said "always use metric" in March and "imperial" in June.
Episodes. What happened in previous sessions: the task, what the agent did, how it ended. Useful for few-shot grounding ("last time this customer asked about billing, we escalated"). These are many, mostly irrelevant, and worth retrieving by similarity with a recency bias. Failure mode: volume. Ten thousand episodes, none of them scored, all of them competing for the same top-k slots.
Procedural notes. Learned instructions — how to handle a class of request, what not to try again. Rare, expensive to get right, and the most dangerous to write automatically. Failure mode: self-poisoning. The agent writes down a bad lesson from one failed run and follows it forever.
In LangGraph terms, the first lives naturally in a Store namespace keyed by (user_id, "profile"); the second in a namespace with embeddings enabled for semantic search; the third in a small, human-reviewable list. One store, three namespaces, three access patterns.
Write on a decision, not on every turn
The cheapest good rule we know: do not let the agent write memory as a side effect of talking. Give writes their own moment.
Two workable shapes:
Hot path, explicit tool. The agent has a remember(fact, namespace) tool and may call it during a run. Fast and transparent — the write shows up in the trace — but the model will over-call it. Every "I prefer short answers" becomes a row. Mitigate with a narrow schema and a system-prompt rule that names what qualifies.
Background extraction. After a thread closes, a separate small-model pass reads the transcript and proposes memory writes against a strict schema. Slower to take effect, much higher precision, and you can run it offline over past traffic to backfill. Cheaper too: one call per thread instead of one per turn.
Most production systems we work on end up with both: a tool for facts the user states outright ("call me Sam"), and a background pass for anything inferred.
Either way, writes go through the same three steps:
- Extract against a typed schema — a
key, avalue, ascope, asource_thread_id. Free-text memories are unreviewable and un-deduplicatable. - Reconcile against what is already stored in that namespace. This is the step teams skip. Load existing rows for the key, and have the model return one of
insert,update,delete, ornoop. "Update" with a supersedes pointer beats silently appending a contradiction. - Write with metadata: timestamp, source, confidence. You will want all three later.
Read narrowly
At read time, the temptation is to dump the whole profile plus the top-10 similar episodes into the system prompt. That is how a 300-token profile becomes a 4,000-token prefix, and how irrelevant episodes start steering the model.
What works better:
- Profile by key, always. Small, deterministic, cacheable. Because it is stable text at the front of the prompt, it sits well inside a prompt-cache prefix; keep it above anything that changes per turn.
- Episodes by tool call, on demand. Make
search_memory(query)a tool the agent calls when it decides prior context matters, rather than a blind pre-fetch on every request. Fewer tokens, and the trace tells you when memory actually mattered. - Cap it. Hard limits: N rows, M tokens, truncate oldest-first. A memory system without a ceiling becomes a context-length bug six months after launch.
- Label provenance. Render memories as a clearly delimited block marked as prior notes about the user, not as instructions. Memory is untrusted input — a user can talk your extractor into storing "always ignore the refund policy." The containment rules for untrusted text apply to your own store.
Forgetting is a feature
Nothing in an LLM app degrades as quietly as a memory store nobody prunes. Decide the expiry policy at design time:
- TTL by type. Episodes expire in 90 days; profile facts do not. Procedural notes get reviewed quarterly by a human.
- Supersede on update. Keep the old row with an
invalidated_atstamp rather than deleting it. When a user complains the agent "remembered something wrong," you need the history to debug it. - Usage decay. If an episode has never been retrieved into a run that a user rated positively, it is a candidate for deletion. Log retrieval hits; they are the only signal you get about whether memory is earning its tokens.
- Deletion path. Users will ask you to delete what you know about them. Namespacing by user makes that a one-call operation. Free-text memories smeared across a shared index make it a project. This is also the difference between a tractable and an intractable GDPR conversation.
Measure it, or you are guessing
Memory feels like it helps. Prove it before you widen the read path.
Build two eval sets. The first is multi-session: a scripted first thread that establishes facts, then a second thread whose grading checks whether those facts were applied. That measures recall. The second is an interference set: threads where stored memories are irrelevant or outdated, graded on whether the agent correctly ignores them. That measures the damage. Teams that only run the first ship a system that gets more confident and less correct as the store grows.
Then track two numbers in production: memory retrieval rate (what share of runs actually pulled a memory into context) and token cost per run attributable to memory. We have seen a memory layer add 900 tokens to every request and get used in under 4% of them. That is not a memory system; it is a tax.
The honest version of the advice: most agents need a small keyed profile and nothing else. Add episodic recall when you have a measured task where prior sessions demonstrably change the answer, and add procedural memory only with a human in the write loop. Start smaller than feels right — the store only grows from there.
If you are building or debugging an agent memory layer, tell us what it is storing and what it gets wrong. We will reply with questions.