LLM Memory: The Boundary Between Context and a Store
LLM memory is the boundary between two stores — the context window, which is rebuilt from the request on every call and dies with it, and a durable memory store that survives the call and is read on later ones. That boundary settles four things: what an agent may assume without being told again, how long a remembered fact lives, which user or session is allowed to see it, and what happens when a…
Short answer: LLM memory is the boundary between two stores — the context window, which is rebuilt from the request on every call and dies with it, and a durable memory store that survives the call and is read on later ones. That boundary settles four things: what an agent may assume without being told again, how long a remembered fact lives, which user or session is allowed to see it, and what happens when a read misses. It is a correctness boundary before it is a retrieval-quality one.
Key takeaways
- Two stores, one decision: a fact is either re-derivable from the request or it has to be written down, and only the second kind earns a row.
- Retention is a promise already made: a time-to-live is set by what the user or an auditor was told, not by the price of storage.
- Scope is the safety boundary: a per-session key forgets a returning user, a per-user key leaks across sessions the moment the key is wrong.
- Recall fails in three distinct ways — a miss, a stale row, a right fact for the wrong caller — and each one needs its own guard.
Our own paid measurement puts this page's phrase plan on a spread of volume rather than one word.
In DataForSEO Google Ads data for the United States over a 12-month window, measured 2026-10-01,
"llm memory" carries 260 searches a month, "tencentdb agent memory" 210, "memory agent" 170,
"context compression" 110, "mcp memory server" 70, "agent memory systems" 50 and "context model
protocol" 40, all recorded in this project's search_volume.json. Those figures describe one layer
of an agent stack, searched for from three directions: as a concept, as a product, and as the
protocol people expect to carry it. This page covers the line that layer draws between the context
a single call is given and the store that outlives it, because most memory bugs are not retrieval
bugs at all — they are bugs at that line.
llm memory: the boundary between a context window and a store
The phrase covers two things that behave nothing alike, and the useful definition is the line between them. The context window is rebuilt from the request on every call: the system prompt, the recent turns, the documents a retrieval step returned, the output of the last tool the agent ran. It is per-invocation by construction, it is paid for again on every turn, and it disappears when the call returns. The memory store is the opposite in every one of those respects — durable rows, written once and read on later calls, shared across processes, and mutable.
Because the two are so different, the boundary is a set of four questions, and each one has an owner. What may be written down, and by whom? How long does a row live before it expires or is superseded? Which caller is allowed to read which row? And what does the system do when a read returns nothing, returns something old, or returns something that belongs to someone else?
The crossing test is short enough to apply at review time. A fact belongs in the store only if it is not re-derivable from the request or from a source the agent can read again, and if losing it would change what the agent does next. A customer identifier the request already carries is context; a preference that appears in no system of record is a row. Applying that test removes most of the transcript a naive implementation would persist, and it keeps the store small enough to stay clean.
The consequence people underestimate is that a memory store is a database with a schema, not a cache with a timeout. A cache may drop an entry silently because the next caller recomputes it; a store may not, because a dropped row changes what the agent believes, and a wrong row changes it in a way that repeats. That is why versioning, provenance and a deletion path are the minimum shape of a store whose contents are treated as true. Memory is not retrieval, which reads a corpus the agent did not write, and it is not compression, which shrinks the current window without changing what exists upstream. The centre page maps all four mechanisms that fill a window and the failure each one owns on context engineering; this page stays on the boundary the memory mechanism draws.
agent memory systems: four jobs, four storages
A memory system is usually sold as one product and is actually four jobs, each with its own storage and its own failure. Naming them separately is what makes a design review possible, because the job a product does badly is the one that ships.
The write decision is the job of deciding, at the end of a run, which facts are worth keeping. Its storage is the run's own trace plus a review step, and its failure is a wrong fact entering the store with the authority of having been written. The store and index is the durable layer itself: rows plus whatever index makes them findable by meaning rather than by key. Its failure is an index that drifts from the rows — a row deleted but still returned, or a vector that no longer matches the text it was built from. The read path is query construction, ranking and, most importantly, the scope filter. Its failure is the one this page returns to below: a confident answer drawn from the wrong caller's rows. The lifecycle is the job nobody owns until an audit asks — time-to-live, supersession, export and deletion. Its failure is a row that outlives the promise made about it.
The input side of the store is worth a sentence of its own, because memory is fed by tool returns rather than by the model's imagination. Whatever the extraction layer hands back is what the write decision can see, so a noisy extraction becomes a noisy memory. The extraction half of that return path — what survives when a document becomes text, and how HTML differs from markdown — is worked through on url to markdown. The four jobs also fail differently, which is why a product that does one of them well can still be the wrong buy for the other three.
memory agent: the recall failures a client has to survive
The read side of memory gets far less attention than the write side, and it is where the visible damage happens, because a read that goes wrong is what the user actually sees. Four failures are distinct enough to need distinct guards, and collapsing them into a single "retrieval quality" metric hides all four.
A miss is a read that returns nothing when a relevant row exists. The common cause is not a bad embedding but a query built from the wrong words: the fact was written as a preference and is looked up as a complaint, or the row is in one language and the question in another. The guard is to write rows in the vocabulary of the questions that will fetch them, and to record, next to each row, the kind of question it is meant to answer. A stale row is a read that returns a fact that was true and is not any more. The guard is not a shorter time-to-live but a supersede relation: the new row names the old one, and the read path prefers the current version while keeping the history for review. A right fact for the wrong caller is the failure with the worst blast radius, because it is indistinguishable from a correct answer to everyone downstream. Its guard is a scope filter applied before ranking rather than after, since a ranking step that can see another user's rows has already leaked them. A conflated fact is two rows quietly merged into one — an address applied to the wrong account — which happens when the write path matches too eagerly on a weak key like a name.
The honest summary of all four is that a memory read must be allowed to answer "I do not know". A store that always returns its nearest neighbour will return something wrong rather than nothing, and a downstream agent will treat it as fact. This is the same discipline that governs retrieval over a corpus, where a rewritten query and a stop condition decide whether a loop converges; that loop's mechanics are the subject of agentic retrieval. Memory adds a requirement the corpus does not have: because the store is mutable and its rows carry scope, every read needs an authorisation decision as well as a relevance decision.
mcp memory server: where user and session scope is enforced
When memory is exposed as a tool on a protocol — an MCP server being the common shape — the scope key becomes the most safety-critical field in the whole system, and the important rule is that the server owns it and the model must not. A client that lets the model pass the user identifier as a tool argument has handed the isolation boundary to the least reliable component in the stack, and the failure is a cross-user read that no amount of prompt discipline prevents.
Four scopes are worth naming separately because each is right for a different fact. Per-call memory is scratch space that lives inside one tool invocation and is discarded with it; it is not really memory and should not be stored. Per-session memory belongs to one conversation thread or working session and is the natural home for the task's own state. Per-user memory crosses sessions and is where durable preferences live, which is exactly why it must be keyed by an identity the caller authenticates rather than one the model supplies. Per-tenant memory is shared by everyone inside an organisation boundary and is the scope where a single wrong row is most expensive. A store that supports only one of these forces the others to be faked with naming conventions, and naming conventions do not survive a helpful model.
The enforcement point matters as much as the scope list. The scope key should be attached by the server from the authenticated caller, the read filter should be part of the query that reaches the store rather than a check applied to its output, and the write path should refuse a row whose scope the caller cannot claim. Where those three live in one component, the boundary is testable; where they are spread across a client, a prompt and a database, a single missing filter is invisible until it leaks. The mechanics of the agent-facing write path — what an agent records and why a write is a decision about the future — belong to the sibling page memory agent; this page's concern is only that the scope travels with the write and is checked on the read.
context model protocol: what a read must carry back
A memory row that carries only text is ungovernable, because the reader has no way to decide whether to trust it. Whatever protocol carries memory between an agent and a store — the Model Context Protocol being the one most implementations reach for — the returned payload needs a small set of fields before any of the boundary above can be enforced. It is worth treating that list as a contract rather than a suggestion, because each missing field removes a guard.
The row needs an identity so it can be superseded or deleted without rewriting a query, a scope so the caller can filter before ranking, and a version with a supersede link so a read can prefer the current fact while keeping the history. It needs provenance — which run wrote it, from what evidence — so a review can ask why this row exists, and a timestamp so staleness is measurable rather than guessed. It needs a confidence or kind label so the read path can weigh a stated preference differently from an inferred one. And it needs an explicit way to say expired, so that "this row is no longer valid" is a value the store can return rather than a row that has silently been removed.
The protocol also constrains the transport. A write should be idempotent, so a retried call does not store the same fact twice under two identities; a read should be bounded and paginated, so a broad query cannot pull an entire user's history into a window billed by the token. And a row's text should never carry a secret, because a memory row is read by the model and therefore by whatever the model is allowed to say.
context compression is not forgetting: two different erasures
These two mechanisms both make text disappear, and confusing them is one of the most common memory bugs, because the fix for one does nothing for the other. Compression is a lossy transform applied inside a single call: a long thread is summarised, a document is shortened, a tool result is truncated before it is placed in the window. The underlying fact still exists upstream — in the transcript, in the store, in the source — and the compression only changes what this call can see. Forgetting is the deletion or expiry of a durable row: the fact is gone, and the next call cannot recover it no matter how it asks.
The practical consequences run in both directions. A team that says "we handle memory by compressing the history" has said nothing about retention, because the store is a separate thing with its own clock; a team that sets a time-to-live and believes it has solved context length has left the window exactly as expensive as it was. Compression has its own honest measure — whether the decision the call had to make survives the transform — and the mechanics of choosing and measuring a compressor are worked through on context compression. Forgetting has a different measure entirely: whether the row is still there when an audit asks, and whether the promise made to the user about its lifetime still holds.
There is one interaction between them that belongs on this page because it is a memory decision. A compressed transcript is often what a later write step reads when it decides what to persist. If compression has discarded the evidence for a fact — the sentence that made it true — then the write has nothing to stand on and records a confident claim with no support. The guard is to make the write decision read from the store of raw events rather than the compressed view: compression stays in the call, and the decision about what to remember should not be made from the copy.
tencentdb agent memory and other open servers: the questions to ask
Memory became a product category quickly, and an open-source memory service such as TencentDB-Agent-Memory — which packages the store and its API so an agent can read and write durable facts without building one — is a reasonable starting point rather than a finished decision. The category is worth evaluating the way any stateful service is evaluated, because the parts that fail are the parts a feature list omits. A fetch layer raises the same shape of question, and the vendor comparison for that layer is drawn out on firecrawl alternative; memory is that comparison with a much longer memory of its own, so to speak, because a fetch can be retried and a wrong row cannot.
The first question is where scope comes from. If the service accepts the user identifier as a parameter, the caller is responsible for isolation, which is fine if the caller is your own service and dangerous if it is a model. Ask for the scope the server can enforce itself. The second is the retention default: what happens to a row nobody touches, and is expiry chosen by you or set for you. A default of permanent is a policy decision disguised as an implementation detail, and it is the one an auditor reads first. The third is export and deletion: can the rows for one user be listed, exported and removed, and is deleting a user a supported operation rather than a migration you write. The fourth is the write review: is every write attributed to a run, and can a bad row be traced and removed without disturbing the rest.
Two more questions separate a service you can operate from one you can only use. Where does the index live, and what happens when it and the rows disagree — a store that ranks on a stale vector returns a fact that was edited away. And how the store behaves under load, because a memory read sits on the critical path of every call. None of these are arguments against a managed memory service; they are the questions whose answers decide whether the boundary this page describes can be held with that service or has to be rebuilt around it.
A memory boundary checklist
Written as a table of decisions, because each row is a place where the default answer is also the wrong one.
| Decision | The default people ship | The boundary-safe answer |
|---|---|---|
| What is written | the whole transcript | a decision or fact that is not re-derivable, with provenance |
| Retention | keep it forever | a time-to-live tied to the promise already made |
| Scope key | the session identifier alone | tenant plus user plus session, set server-side |
| Read filter | rank by similarity | filter by scope first, then rank |
| Stale rows | overwrite in place | version and supersede, keep the prior row |
| Deletion | leave the row where it is | delete by scope, and list the keys before you do |
The order of the rows is the order of the risk. Getting the write decision right removes most of the rows that would otherwise need governing; getting retention wrong outlives everyone who made the choice; getting scope wrong is the only one of the six that can be a breach rather than a bug.
Where SmartGate fits
Memory is not a feature of a gateway, and this page will not pretend otherwise — the store is a database and the boundary above is a property of your application. What a gateway does own is the part of the boundary that lives on the call path, and three of its properties push the memory design in the right direction. A per-key token budget makes the context-versus-store split economic rather than philosophical, because the window is billed and the store is not: once the budget is a number, "what do we send again on every turn" becomes a question with a price. A retention window on the call record gives episodic memory a floor, since the audit rows and the episodic facts are usually the same events seen from two sides. And a per-key identity is exactly the authenticated scope key a memory server wants, which is what makes the isolation boundary enforceable rather than aspirational.
The plan table is a set of operational limits rather than a feature list: monthly token caps of 2M, 20M, 100M and 200M+; MCP requests per minute per key of 120, 300, 600 and 1200; audit-log retention of 7, 30, 90 or 180 days; and 2, 10, 30 or 9999 keys per team. Treat the pricing page as the authoritative table for all four figures and check it before committing to a retention window or quoting a limit, because the numbers above are the ones in force on 2026-10-01 and they change with the plan.
Frequently Asked Questions
Limitations
This page describes a boundary and the decisions at it; it does not benchmark any memory product or rank open-source servers against each other. Named projects appear because a measured search phrase or a source citation carries their name, and the sentence attached to each one points at the question its own documentation answers rather than at a verdict this page is not in a position to give.
The failure modes named above are drawn from how the mechanisms behave, not from a measured incident count, and the guards are engineering patterns rather than guarantees. A scope filter, a supersede relation and a review step each reduce a class of error; none of them makes a memory store safe to leave unsupervised, and none of them substitutes for testing the read path with two real caller identities.
The plan figures are operational limits read on one date and re-verified against the live pricing page on 2026-10-01. They are not a feature comparison and they change with the plan. This page carries no code excerpt on purpose; the reason is recorded in the Method note below.
Sources
- The Model Context Protocol specification — modelcontextprotocol.io/specification, for the tool and resource shapes a memory server is reached through, and the client-server split that decides where scope can be enforced.
- The reference memory server in the Model Context Protocol servers collection — github.com/modelcontextprotocol/servers, as the shape of a memory service exposed as an MCP tool.
- TencentDB-Agent-Memory — github.com/TencentCloud/TencentDB-Agent-Memory, as an open-source agent memory service, cited for the questions its documentation answers rather than for a ranking.
- Mem0 documentation — docs.mem0.ai, as a managed memory layer with its own account of the write and read paths.
- LangGraph memory documentation — docs.langchain.com, for the short-term versus long-term memory split and its scoping model.
- The MemGPT paper — arxiv.org/abs/2310.08560, as the origin of the paged-memory framing that most agent memory designs borrow from.
- A survey of memory mechanisms in language-model agents — arxiv.org/abs/2404.13501, for the write-manage-read decomposition this page's four jobs follow.
- Demand figures quoted in this page are our own measurements: DataForSEO Google Ads, United
States, 12-month window, measured 2026-10-01, recorded in this project's
search_volume.jsonandresearch_brief.md. - Product behaviour and the plan table: read from the product source at the revision recorded in
this project's
pipeline_results.json, read-only, with the plan figures re-verified against the live/pricingpage on 2026-10-01. That page is authoritative and is where a commitment should start.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher pinned 0 of 7 sections for this page (0 abstention(s), 7 no-slice verdict(s)): rule A found no unique symbol in the scanned repository for any of the seven section phrases, and every recorded candidate list was empty rather than ambiguous, because this lane's vocabulary — memory, context, agent — collides with generic module, type and helper names across a product codebase. A pinned generic would have given the page the shape of a verified article with none of the substance, so the house rule for an unpinned section applies and every section above is written from sources.
Product claims were read from the product source at the revision the slice run recorded in this
project's pipeline_results.json, read-only, and the plan figures were re-verified against the live
pricing page on 2026-10-01. The section keyword quoted above each heading comes from this project's
own paid measurement run, not from a third-party tool. No code, batch fingerprints, auction data or
internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.