SmartGate

Memory Agent: When to Write, Read, and Forget

A memory agent is the interface between one agent run and everything that should outlive it. A write is triggered by a decision, a correction or a stated preference and refused when the unit is a cache entry, a derived value or a guess; a read returns one stated unit of recall inside a mandatory scope and a score floor; contradictions and expiry are settled by a policy written before they happen.

Short answer: A memory agent is the interface between one agent run and everything that should outlive it. A write is triggered by a decision, a correction or a stated preference and refused when the unit is a cache entry, a derived value or a guess; a read returns one stated unit of recall inside a mandatory scope and a score floor; contradictions and expiry are settled by a policy written before they happen. The storage engine is a separate question, and this page is about the four decisions in that interface.

Key takeaways

  • Write on a trigger, not on a schedule. A decision, a correction, a stated preference or a confirmed identity fact earns a write; a tool result a later step will re-fetch does not.
  • A read returns a unit, so the unit is a design choice. A raw turn, a rolling summary, an atomic fact and an entity profile carry different write costs and different review burdens.
  • Conflicts and expiry are one policy, not two. Decide the rule — newest wins, source of truth wins, or a human decides — before the contradiction arrives, and give every unit an owner and a lifetime.
  • Every read and write is an audit event. Memory is the second place a secret can leak from, so the scope lives in the filter and the boundary is a property of the interface.
  • Measure the read that found nothing. A store that answers most reads needs a better query, not more vectors.

memory agent: the four decisions behind one interface

A memory agent is not a database and not an embedding model. It is the component that sits between a run and whatever should survive it, and it earns that name by making four decisions explicit instead of leaving them to a framework default. The first is the write trigger: which kind of event is allowed to create a memory. The second is the unit of recall: what one read hands back — a raw turn, a rolling summary, an atomic fact, or a profile. The third is conflict and expiry: which of two contradictory units wins, and when a unit stops being true. The fourth is the boundary: which scope a write lands in, who may read it, and which audit row records both.

Everything else people argue about — the distance metric, the index type, one store or two — is downstream of those four, and none of it is decided by them. A team that picks a vector store first and writes the trigger rules last ships a store that grows monotonically and answers questions nobody asked. A team that writes the four decisions down first can replace the storage engine without rewriting the policy, because the policy never mentioned it.

The interface is small enough to state in two calls. A write takes a candidate unit, a scope and a source; it either lands or is refused, and the refusal is a decision the caller can log. A read takes a query, a scope, a score floor and a cap; it returns at most k units, each with the score that admitted it and the provenance that produced it. Those two calls are the contract, and the engineering worth doing is in the arguments rather than in the transport between them.

It helps to see where this sits. The window an agent reasons over is filled by four mechanisms — retrieval, compression, memory and tool return — and memory is the only one whose content outlives the process that created it. Which mechanism owns which failure is the map drawn on context engineering; this page takes the memory mechanism alone and works the interface. The loop and the compressor are worked through on their own pages, and this one points at them rather than restating them.

agent memory systems: when a write is triggered, and when it is refused

An agent memory system is only as good as its refusal rate. Writing is cheap at the moment of the call and expensive forever after: every stored unit is context a future read may pay for, and a wrong unit is a claim that later gains authority simply by having been retrieved before. The trigger is therefore a policy, and the useful version of it names four events that justify a write and a longer list that does not.

Write when a decision is made — the choice, the reason, and the alternative that was rejected. Write when a correction arrives, because a user who says the balance is on the other account has just told you which of your memories is wrong, and that correction is the most valuable unit in the store. Write when a preference is stated in the user's own terms, since preferences are stable and are exactly what a later run cannot recompute. Write when an identity fact is confirmed — an account id, a repository, a timezone — provided the confirmation is direct rather than inferred.

Refuse the write in the cases that fill most stores by accident. A tool result a later step will re-fetch is a cache entry, not a memory. A derived value such as a total, an embedding or a formatted report can be recomputed, so storing it only creates a second place for it to go stale. Transient run state belongs to the run and dies with it. A model's own plausible-sounding assertion is a hypothesis, not an observation, and writing it converts a hypothesis into a fact with a retrieval history. Credentials must never be written, because a memory payload is precisely what the read path filters on and returns. And a record that already lives in a system of record should be referenced rather than copied, or the copy becomes a silent authority on the next read.

Two mechanical properties follow from the trigger. The first is idempotency: a retried write must not create a second copy of the same unit, so the writer supplies a stable key instead of letting the store mint one per attempt. The second is attribution: every unit records which run produced it and from what evidence, so a later contradiction can be traced to a source rather than argued about. Where a unit is fed by a fetched page, the quality of that page decides the quality of the memory, and the path from HTML to clean text is the subject of URL to markdown. Refusing a write is also the cheapest thing an agent can do: a unit never stored is a unit never embedded, never indexed and never retrieved by mistake.

llm memory: what one read returns, and at what granularity

LLM memory is usually described as a storage tier, but the decision that changes an agent's behaviour is the unit of recall. There are four honest units, and they are not interchangeable.

A raw turn — the literal exchange, stored as observed — is the cheapest to write and the most expensive to read: a read that returns five turns spends five turns of context to deliver one fact. A rolling summary of a session is compact and pleasant to read, and it loses the exact wording a correction or a commitment depends on. An atomic fact — one claim, one subject, one provenance — is precise, filterable and cheap to read, and it multiplies the number of writes, which moves the cost back to the trigger policy above. An entity profile aggregates everything known about one subject and is the right unit when the question is what we know about this customer, and the wrong one when the query is narrow, because a profile read hands back facts the caller did not ask for.

The unit also decides how a read can be judged. Facts admit a score floor, a metadata filter and a cap; summaries admit only a similarity score, which is a weaker basis for admitting text into a window. That is the mechanism behind an observation that surprises people: the more mature the system, the smaller the unit. A small unit is the only one you can refuse cheaply.

Whichever unit you choose, a read takes the same four arguments and all four are policy. Scope says whose memory this is; an unscoped read is a data-leak path, so the scope is mandatory rather than a convenience. Score floor decides which candidates are eligible at all. Cap bounds what a single read costs in context. Reranking, when it is offered, buys a second model call on every read and belongs only where a wrong unit is expensive. A read must also be able to return nothing above the floor, and that empty result is an answer rather than an error. The read side and its scoring are worked through in more depth on LLM memory; what matters at the interface is that recall is a thresholded, scoped, capped operation and not a search box.

agentic retrieval: the read as a step in a loop, not a lookup

The older shape of retrieval was one lookup in front of one prompt. The shape agents actually run is a loop: the agent forms a query from its current goal, reads, inspects what came back, and decides whether to read again with a sharper query. That change moves three properties out of the application and into the interface.

First, a read must be cheap enough to repeat, because the loop may run it several times inside one task. A memory read that costs a model call to rerank, or that scans an unscoped index, turns an iterative loop into a bill. Second, a read must return its own evidence — the score, the provenance and the age of each unit — so the caller can decide whether a second query is worth issuing or whether the store simply does not know. An interface that returns bare text forces the model to guess at relevance, and that guess is where hallucinated confidence enters a run. Third, a read that returns nothing must be distinguishable from a read that failed: an empty result above the floor is information, a transport error is not, and collapsing the two into one empty string is how a silent recall loss ships to production.

The loop also changes what should be written. When retrieval is iterative, the agent learns things while searching — a refined query that worked, a source that turned out to be authoritative — and those observations are memory candidates with the same attribution requirement as any other. The mechanics of the loop itself, from query rewriting to the stopping condition, belong to agentic retrieval. The requirement this page adds is narrower and easy to state: a read must be repeatable, attributable and honest about emptiness, because it will be called more than once.

context compression: keeping the store small enough to be worth reading

Compression is usually discussed as a prompt-time trick, and it is also the cheapest form of memory hygiene. The reason is arithmetic: a read returns the top units above a floor, so a store full of near-duplicates spends its cap repeating one fact five ways while the genuinely new unit never makes the cut. Compressing the store is not cosmetic; it is what keeps the read useful.

Three operations do most of that work, and none of them needs a model present at read time. Deduplicate on write: two units describing the same subject with a similarity above a threshold are one unit, and the policy decides whether the newer replaces the older or is merged into it. Merge deliberately when several partial observations add up to one fact, keeping the union of their provenance so the merged unit stays traceable. Expire by rule: a unit with a validity window — a price, a temporary role, a project that has ended — should carry that window at write time, because expiry decided later is a migration and expiry decided at write time is a field.

The dangerous default is newest-wins applied silently. It is correct for a correction, where a user has explicitly overridden an earlier value, and wrong for a disagreement, where two sources are simply saying different things and the honest state is that they conflict, with each provenance attached. A policy that cannot represent a conflict resolves it by accident, and the accident stays invisible until the read is wrong in production. Deleting is the other half of expiry and deserves the same treatment: a unit removed for a correction or a privacy request should leave a tombstone, so a later read can tell that something was removed rather than that nothing was ever known. Where the store is fed by fetched material, the compression that matters most happens before the write, and the fetch layer is where that noise enters — choosing it deliberately, including a self-hosted path, is the comparison on Firecrawl alternative. What a compressor itself drops, and how to measure the loss, is the subject of context compression.

tencentdb agent memory: what a shared, team-level store changes

A team-level memory product — the shape the TencentDB Agent Memory project popularised, where one store is shared across agents and people rather than scoped to a single run — is worth reading as a set of interface changes rather than as a competitor to rank. When memory is shared, the four decisions stop being local to one agent.

The write trigger gains an authorisation question: a unit written by one agent is now readable by others, so is this worth keeping is joined by is this mine to publish. The unit of recall gains an audience: a fact exact enough for one engineer may be a personnel detail to another, so the unit needs a sharing level and not only a scope. Conflict stops being internal, because two agents will disagree about the same customer and the resolution rule now has to be one the whole team accepts. And the boundary becomes the main engineering: who may read a shared scope, what a shared write records about its author, and how a person disputes a unit an agent attributed to them.

None of that is exotic, and none of it is solved by a larger index. Shared memory is a governance problem wearing a storage costume, which is why the honest adoption path is to model a team store as a set of scopes with an owner each and a review path for the semantic units, rather than as one pool everything may write to. Treating a team memory as a single unscoped bucket is the same error as an unscoped read, one level up. The payoff is real when the model fits: a team store is what lets one agent's correction reach another agent's next read, and that is the only way memory earns its cost across a fleet rather than inside a single process.

mcp memory server: the surface where scope, audit and privacy meet

Most agents reach memory through a tool rather than a library, and the tool surface is where three of the four decisions stop being conventions and become enforced. A memory server is best modelled as one tool with an action parameter — add, search, delete — instead of a family of endpoints, because one declaration is one thing to authorise, one thing to document and one thing to audit.

The scope is enforced in the read filter rather than in the index, so a read that names no scope is refused before the store is touched. The audit row is what makes the store operable: caller, transport, action, scope, the number of units returned and the outcome, written once per call, so that a memory read and every other governed call produce the same shape of record. That row is also what answers the questions a team actually asks — which agent wrote this, when, and on what evidence — and it is the reason attribution belongs at write time instead of being reconstructed later.

The privacy boundary has two halves and only one of them is technical. The technical half treats the store as public within its scope: a payload is what the read path filters on and returns, so a credential written into one is a credential a later read may surface, and the sound rule is to keep secrets out of the store entirely rather than to mask them on the way out. The other half is retention. Memory inherits the same obligation as any other record, and the system remembers it is not a reason a unit may be kept past the retention policy that governs the log it came from. Making retention a field on the unit, and the boundary a property of the interface, is what keeps the answer to a question like what does this agent know about me a query rather than an investigation.

Where SmartGate fits

SmartGate exposes memory as one of its algorithm primitives behind the same MCP endpoint as every other tool, which makes it a concrete instance of the interface above rather than a library to embed. A write carries a scope and a source; a read carries a query, a scope and a floor; and both produce an audit row in the same table as a fetch or a compression call, so the boundary and the attribution are properties of the gateway instead of properties of each agent that uses it. One write path, one read path and one record is a smaller surface to reason about than a framework class embedded in every service.

The operational limits are plan limits, and they are worth reading as interface constraints rather than as a price sheet. Monthly token caps run across four tiers at 2M, 20M, 100M and 200M+; MCP requests per key per minute are 120, 300, 600 and 1200; audit-log retention is 7, 30, 90 or 180 days; and a team may hold 2, 10, 30 or unlimited keys. Two of those bear directly on this page. Retention sets the floor on how long a unit's provenance stays auditable at all, so a memory whose evidence must be defensible longer than the plan's window belongs somewhere else as well. And the per-key request rate sets how often an iterative read loop can afford to re-query, which is why the read-path design above keeps one read cheap. The table is authoritative — read it at /pricing — because a compliance or capacity decision has to rest on the current numbers rather than on this paragraph.

Frequently Asked Questions

Limitations

This page describes an interface and the policies behind it; it does not benchmark a store, rank a vector database, or claim that any named project implements the decisions well. The unit sizes, thresholds and caps above are the shapes to reason about, not recommended constants, because the correct floor for one workload is a rate limit for another.

The four decisions are necessary and not sufficient. They say nothing about embedding quality, index recall or the latency of a read, all of which are properties of the engine behind the interface and are outside its scope. A perfectly specified policy on top of a weak index still returns the wrong units.

A conflict policy that can represent disagreement is not the same as one that resolves it correctly. This page argues for recording both units with their provenance; it does not claim that any automatic rule is right in general, and a store that surfaces a conflict to a human has to have somewhere for that decision to be made.

The retention and scope discussion assumes the caller supplies an honest scope. Where a host passes one shared identity across several tenants, every isolation property described here is defeated by the caller rather than by the interface, and no amount of filtering inside the store repairs that.

The product limits quoted above are operational limits drawn from the plan table, not a feature comparison. Caps, per-key rates, retention windows and key counts change with the plan, and a compliance decision has to read the current table rather than this page.

This page carries no code excerpt, and the reason is recorded in the Method note below: the slice matcher pinned no section. The consequence is honest and worth stating plainly — there is no quoted implementation detail and no line-numbered claim anywhere above.

Sources

Method note

This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher pinned 0 of 7 sections for this page (0 abstention(s), 7 no-slice verdict(s)): rule A found no unique symbol in the scanned repository for any of the seven section keywords, because this lane's vocabulary — memory, agent, retrieval, compression — collides with generic type and module names across a codebase, and the remote symbol-candidate fallback returned only score-ranked, non-unique names (the same Memory.add family for three different phrases). A pinned generic would have given the page the shape of a verified article with none of the substance, so every section above is written from sources.

Product claims were read from the product source at the revision the slice run recorded in this project's pipeline_results.json, read-only, and the plan figures were re-verified against the live pricing page on 2026-10-01. The section keyword quoted above each heading comes from this project's own paid measurement run rather than from a third-party tool. No code, batch fingerprints, auction data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.