Agent Memory for Agentic Systems
Building anomaly detection and churn prediction agents that learn, with mem0, Hindsight, and cognee
TL;DR
Stateless agents re-read the entire conversation every turn, so cost grows quadratically with turns — durable memory is what breaks that curve.
Memory needs distinct layers with different write paths and read costs: raw events, consolidated evidence-backed beliefs, and precomputed answers to questions asked every session.
Retrieval should run against a hard token budget, not "include everything that might be relevant," and each session should distill into durable knowledge only under an explicit gate.
Evaluated against three open-source systems — mem0, Hindsight, and cognee — for two production workloads: an anomaly detection agent that should stop re-investigating closed patterns, and a churn agent that should stop repeating discount plays that already failed.
The assignment
The starting point. Two agentic workloads are already running in production and both are underperforming for the same reason. The anomaly detection agent investigates the same recurring signal every night, because nothing tells it that a human already looked at this pattern last Tuesday and closed it as expected seasonal behaviour. The churn prediction agent recommends a discount play to an account that was given the same discount play six weeks ago, which did not work, because nothing tells it that either. Both agents are competent within a single conversation and amnesiac across conversations.
What success requires. The agents need durable memory that survives session boundaries, model upgrades, and restarts. That memory has to be auditable, because both workloads drive actions that a human will later be asked to justify. It has to be governed, because churn memory contains customer data and anomaly memory contains infrastructure detail. And it has to reduce cost rather than add to it, because the naive approach — keep more context around — makes every turn more expensive than the last.
The approach. Separate memory into layers with distinct write paths and read costs. Retrieve with a hard token budget rather than a hope. Consolidate raw events into evidence-backed beliefs. Precompute the answers to questions asked every single session so they become a stored read instead of a retrieval plus a model call. Distil each session into durable knowledge at the end, under a gate, so only confirmed lessons persist. Then instrument the whole thing so that tokens per resolved task is a number someone owns.
What it delivers. Cost per turn that tracks the difficulty of the question instead of the length of the transcript. An anomaly agent that suppresses with a reason rather than re-deriving. A churn agent that recommends with a record rather than a guess. And an evidence trail that makes both defensible in a review.
This article works through that design against three open-source memory systems — mem0, Hindsight, and cognee — and is careful to distinguish what each project documents from what an architect still has to decide.
Why a stateless agent is expensive
The economics are simple enough to write down, and they are worse than they look.
In a stateless loop the model sees the whole conversation every turn. If turn $i$ adds \(d_i\) tokens of new material and the system prompt plus tool schemas cost a fixed $S$, then the input tokens billed at turn $n$ are:
$$T_n = S + \sum_{i=1}^{n} d_i$$
and the total billed across an $N$-turn session is:
$$T_{\text{total}} = \sum_{n=1}^{N} \left( S + \sum_{i=1}^{n} d_i \right) = N \cdot S + \sum_{i=1}^{N} (N - i + 1), d_i$$
The second term is the problem. Material introduced early is re-billed once per remaining turn. A document pasted at turn 2 of a 30-turn investigation is paid for 29 times. The growth is quadratic in turns, which is why long agent sessions have a cost curve that surprises people who budgeted linearly.
With bounded recall the shape changes. If recall is capped at $B$ tokens regardless of history:
$$T^{\text{mem}}{\text{total}} = \sum{n=1}^{N} \left( S + B + d_n \right) = N(S + B) + \sum_{i=1}^{N} d_i$$
Now it is linear in $N$. The saving for a long session is the difference between the two, and it is dominated by the quadratic term you removed.
A worked example. Take a 30-turn anomaly investigation with \(S = 3{,}000\) tokens of system prompt and tool schemas, and an average \(d_i = 1{,}200\) tokens per turn.
Stateless: \(T_{\text{total}} = 30 \times 3{,}000 + \sum_{i=1}^{30}(31-i) \times 1{,}200 = 90{,}000 + 1{,}200 \times 465 = 90{,}000 + 558{,}000 = 648{,}000\) input tokens.
Bounded recall at \(B = 4{,}000\): \(T^{\text{mem}}_{\text{total}} = 30 \times 7{,}000 + 30 \times 1{,}200 = 210{,}000 + 36{,}000 = 246{,}000\) input tokens.
That is roughly a 62 percent reduction on a single session, and the gap widens with every additional turn. It is also the reason the first lever in the next section matters more than the other three combined.
Two caveats before anyone puts that number in a business case. First, memory systems make their own model calls — extraction on write, sometimes reranking on read — so the saving is not free and has to be measured end to end, not just on the agent's own prompts. Second, prompt caching changes the arithmetic for the stable prefix, and any honest comparison has to account for it. Neither caveat changes the direction of the result; both change its size.
What "agent memory" actually means
"Add memory to the agent" is not one decision. It is at least five, and conflating them is the most common way these projects go wrong.
| Layer | Holds | Written when | Read cost | Failure mode if missing |
|---|---|---|---|---|
| Working | Current turn, scratchpad, tool results in flight | Continuously, in process | Free, already in context | None — this layer always exists |
| Episodic | What happened and when: runs, alerts, interactions | On event, usually asynchronously | Retrieval + ranking | Agent cannot say "we saw this before" |
| Semantic | Entities, relationships, facts, time series | On extraction from episodic material | Graph traversal or vector search | Agent cannot connect two things that are connected |
| Consolidated belief | What we now think is true, with supporting evidence | On consolidation, batched | Retrieval, but far fewer items | Agent re-derives the same conclusion forever |
| Procedural | Learned plays, rules, standing answers | On distillation, gated | Stored read — cheapest of all | Agent never improves, only remembers |
The last two layers are where the token savings live, because they are read as stored text rather than reconstructed from raw material. Hindsight is explicit about this: reading a mental model is a database read with no retrieval and no LLM call. That is the difference between paying to remember and having already remembered.
Most teams implement the episodic layer, call it done, and are then puzzled that costs went up. Episodic memory alone adds a retrieval bill without removing a reasoning bill.
The three systems, honestly compared
All three are legitimate open-source projects with published research behind them. They are not interchangeable, and the differences are architectural rather than cosmetic.
| mem0 | Hindsight | cognee | |
|---|---|---|---|
| Licence | Apache-2.0 | MIT | Apache-2.0 |
| Paper | arXiv:2504.19413 | arXiv:2512.12818 | arXiv:2505.24478 |
| Core metaphor | A memory layer for agents | Agents that learn, not just remember | An AI memory platform built on a knowledge graph |
| Primary operations | add, search |
retain, recall, reflect |
remember, recall, improve, forget |
| Retrieval | Multi-signal: semantic, BM25, entity — scored in parallel and fused | Four strategies in parallel: semantic, keyword, graph, temporal — fused by reciprocal rank, cross-encoder reranked, trimmed to a token limit | Hybrid selection across graph, vector, and code context, with automatic or chosen strategy |
| Isolation model | User / session / agent scoping | Banks — isolated stores with documented no-cross-bank-leakage | Datasets with a permissions model |
| Distinctive asset | Single-pass ADD-only extraction; agent-generated facts first-class | Observations and mental models — consolidated beliefs and standing answers | Ontology support and custom data models; code becomes a symbol graph |
| Runs without an LLM key | No | No | Yes — local GLiNER extraction and local embeddings |
| Built-in secret/PII scanning | Not documented as a built-in | Yes — Memory Defense, opt-in per bank, 45 patterns, redact or block | Not documented as a built-in |
mem0
mem0's April 2026 memory algorithm reports the following on public long-memory benchmarks, using single-pass retrieval at a top_200 budget:
| Benchmark | Score | Tokens | Latency |
|---|---|---|---|
| LoCoMo | 92.5 | ~7.0K | ~0.88s |
| LongMemEval | 94.4 | ~6.8K | ~1.09s |
| BEAM (1M) | 64.1 | ~6.7K | ~1.00s |
| BEAM (10M) | 48.6 | ~6.9K | ~1.05s |
The token column is the interesting one for this article. Accuracy in the nineties at a roughly 7K token working set is the claim that makes memory look like a cost reduction rather than a cost addition.
Read the caveat that ships with those numbers, though: the project states the scores reflect the managed platform, which includes proprietary optimisations, and that the open-source package is directionally similar but not identical. mem0 has open-sourced its evaluation framework at mem0ai/memory-benchmarks, which is the right response and means you can re-run this against your own data rather than taking the table on trust. Do that before quoting the numbers internally.
Architecturally, the most consequential change in that release is single-pass ADD-only extraction: one LLM call per write, with no UPDATE or DELETE phase. Memories accumulate and nothing is overwritten. That is cheaper and more auditable, and it moves the burden of reconciling contradictions onto retrieval, which is where temporal reasoning and entity linking earn their place.
from mem0 import Memory
memory = Memory()
# Agent-generated facts are first-class, not just user utterances.
memory.add(
[
{"role": "user", "content": "Latency on checkout-api spiked at 02:10."},
{
"role": "assistant",
"content": (
"Confirmed: cause was the nightly reindex job, not a regression. "
"Benign, recurs weekly."
),
},
],
user_id="sre-oncall",
)
hits = memory.search(
"checkout-api latency spike overnight",
filters={"user_id": "sre-oncall"},
top_k=3,
)
The [nlp] extra adds the BM25 and entity-extraction path that makes retrieval multi-signal rather than purely vector-based. Install it; pure vector recall is exactly where this class of system disappoints on entity-heavy queries like hostnames and account IDs.
Hindsight
Hindsight is the one with the strongest position on the layers that reduce tokens, and it has the most unusual validation story: the project reports state-of-the-art results on LongMemEval, independently reproduced by the Virginia Tech Sanghani Center and by The Washington Post. Independent reproduction is rare in this space and worth more than a self-reported leaderboard.
Its three operations map cleanly onto the layered model:
retain— an LLM extracts key facts, temporal data, entities, and relationships, then normalisation resolves them into canonical entities, time series, and search indexes.recall— four retrieval strategies run in parallel — semantic, keyword/BM25, graph (entity, temporal, and causal links), and temporal range filtering. Results are merged, ordered by reciprocal rank fusion, cross-encoder reranked, and trimmed to fit within the token limit. That last clause is the whole token-control argument in one phrase: the budget is an input, not an outcome.reflect— deeper analysis over what has been retained, forming new connections that no single retain could have produced.
Two constructs matter more than the rest:
Observations are consolidated beliefs. Related facts are deduplicated into a single belief that carries its supporting evidence, exact quotes, and a proof count. Beliefs are refined rather than overwritten, so the trail of how a conclusion strengthened remains intact. For an anomaly agent, an observation is "this signature on this host is a reindex artefact" with eleven pieces of evidence behind it, instead of eleven separate alert records.
Mental models are the standing answer to a recurring question about a bank. The project's own framing is the key sentence: reading one is a database read, with no retrieval and no LLM call. Presented as knowledge pages, they are living markdown documents that update as understanding changes. An agent that boots by reading its mental models starts the session already knowing things.
Banks are the isolation unit — one memory store per user, agent, or project, with documented strict isolation and no cross-bank leakage. Banks also carry background context and disposition traits such as skepticism, literalism, and empathy, which shape how aggressively the system forms beliefs from thin evidence.
Memory Defense is the governance feature that matters for the churn use case: an opt-in per-bank policy that scans every retain for secrets and PII against 45 patterns, then either redacts — producing markers like [REDACTED:github_token] — or blocks the write entirely. Scanning at write time rather than read time is the correct choice, because it means the sensitive value never lands in storage at all.
Operationally it runs on PostgreSQL with pgvector, or Oracle AI Database 23ai. It exposes Prometheus metrics for LLM calls, tokens, and latency, which is exactly the instrumentation this article argues you need. It ships an MCP endpoint per bank, and an LLM wrapper that retrofits memory onto an existing client in two lines:
from openai import OpenAI
from hindsight import wrap_openai
client = wrap_openai(
OpenAI(),
bank_id="churn-emea",
hindsight_api_url="http://localhost:8888",
)
cognee
cognee's distinguishing property for enterprise work is that memory is a knowledge graph with first-class ontology support, and that it can run with no LLM key at all — local GLiNER extraction and local embedding models, downloaded on first use. For a regulated environment evaluating memory before any data leaves the perimeter, that is a genuinely different starting position from the other two.
import asyncio
import cognee
async def main():
await cognee.remember(
"Account 88213 moved from weekly to monthly logins in Q3 "
"after the admin left.",
dataset_name="churn_emea",
)
for result in await cognee.recall(
"Why did account 88213 reduce usage?",
datasets=["churn_emea"],
):
print(result)
asyncio.run(main())
The four verbs — remember, recall, improve, forget — include the one the others under-emphasise. forget removes a specific item or dataset, and in a churn system with customer data subject to erasure requests, a documented deletion primitive is not a nice-to-have.
Session distillation curates accepted lessons from a session into permanent memory. That is lever four in the next section, implemented as a product feature, and the word accepted is doing important work: it implies a gate rather than automatic promotion.
cognee's own BEAM results are reported with unusual candour: 0.79 at 100K tokens using fixed hybrid retrieval, and 0.67 at 10M tokens described explicitly as an exploratory result with question-type routing selected on the same question set. The project notes the two settings use different conversations, ingestion models, and retrieval-selection procedures, and asks readers to read the methodology before comparing against other systems. Take that request seriously — it applies equally to every other number in this article.
Two practical notes. Running the entire memory layer on a single PostgreSQL instance is available but currently released as a demo feature, with the production-ready version licensed; do not build a production topology on the demo path without checking that. And cognee supports importing memory from mem0, Letta, Zep, or Graphiti via the COGX exchange format, which materially lowers the cost of being wrong about your first choice.
The four levers that actually reduce tokens
Lever 1 — Bounded recall
Retrieve from several indexes in parallel, fuse the rankings, rerank, then trim to a hard budget. Both Hindsight and mem0 describe parallel multi-signal retrieval; Hindsight additionally describes explicit trimming to a token limit.
Rank fusion matters here because the signals are not comparable on score. Semantic similarity, BM25 relevance, and graph proximity produce numbers on different scales, and normalising them is fragile. Reciprocal rank fusion sidesteps the problem by using position rather than score. For a document $d$ appearing at rank \(r_s(d)\) in each retriever $s$ within the set $S$:
$$\mathrm{RRF}(d) = \sum_{s \in S} \frac{1}{k + r_s(d)}$$
with $k$ conventionally 60. The constant damps the influence of top ranks so that one retriever cannot dominate the fused list on its own, and a document that places respectably across several retrievers outranks one that places first in exactly one.
Work a small case. Three retrievers return memory items, and item \(M_3\) appears in all three at middling rank while item \(M_1\) tops the semantic list only:
| Item | Semantic rank | BM25 rank | Graph rank | RRF score |
|---|---|---|---|---|
| \(M_1\) | 1 | — | — | \(\tfrac{1}{61} = 0.0164\) |
| \(M_3\) | 4 | 3 | 2 | \(\tfrac{1}{64} + \tfrac{1}{63} + \tfrac{1}{62} = 0.0476\) |
| \(M_7\) | — | 1 | 5 | \(\tfrac{1}{61} + \tfrac{1}{65} = 0.0318\) |
\(M_3\) wins, which is the behaviour you want: corroboration across independent signals beats a single confident vote. This is also precisely why a memory system that only does vector search underperforms on queries containing identifiers — an account number is a BM25 and entity signal, not a semantic one.
Then trim. Fill the budget in fused-rank order and stop. If you never hit the ceiling, the ceiling is set too high and is not doing any work.
Lever 2 — Consolidation
Collapse many raw events into one belief carrying its evidence. Fifty alerts on one host become a single standing finding with a proof count. This shrinks the corpus that retrieval searches, which compounds with lever one: a smaller, denser corpus means the same token budget carries more distinct information.
Keep the evidence attached. A belief with no quotes and no proof count is an assertion, and the first reviewer to ask "how do we know that?" will end the project's credibility if there is no answer.
Lever 3 — Standing answers
Some questions are asked at the start of every single session. What is normal for this service? What is this account's history with us? Precompute them. Hindsight's mental models and knowledge pages are this lever as a product feature, and the payoff is that the read costs nothing but a database fetch.
This is the lever most teams skip and the one with the best ratio of effort to return, because it removes work rather than making work cheaper.
Lever 4 — Distillation
At session end, write the lesson and discard the noise, behind an acceptance gate. cognee's session distillation is the documented instance. Ungated distillation is how a memory system slowly poisons itself, so the gate is not optional.
Use case A — Anomaly detection
An anomaly detection agent's memory problem is suppression with a reason. The detector fires; the agent must decide whether this is genuinely new or the fourth occurrence of something already understood. Without memory it investigates from zero every time, which is both expensive and corrosive to trust, because an agent that raises the same resolved issue repeatedly gets ignored.
What to store
| Memory | Content | Layer | Written by |
|---|---|---|---|
| Baseline | What normal looks like for this entity at this hour and season | Consolidated belief | Batch consolidation over historical episodes |
| Fired alerts | Already raised, already open, already linked to an incident | Episodic | Detector, on fire |
| Dismissals | Which alerts a human closed, and the reason given | Consolidated belief | Human action, promoted after review |
| Causal links | Which upstream change last explained this signature | Semantic | Extraction from resolved incidents |
| Runbook outcomes | Which diagnostic steps produced signal on this signature | Procedural | Distillation at investigation close |
The flow
alert fires
→ recall(signature, entity, window) # bounded, ~4K tokens
→ read standing answer: "normal for this service" # stored read, 0 LLM calls
→ is this covered by an existing observation?
yes → suppress, cite the observation and its proof count
no → investigate, then retain the finding
→ on close: distil the lesson behind a review gate
The suppression branch is the one that pays. It costs a bounded recall plus a stored read, and it replaces a full investigation. The economics of the whole workload are determined by what fraction of alerts take that branch.
The honest difficulty
Suppression is a safety-relevant decision. An agent that suppresses a genuine incident because a superficially similar one was benign has caused harm that a noisy agent would not have. Three controls follow from that, and none are optional:
Suppress only against observations with a proof count above a threshold you set deliberately. One prior dismissal is not evidence.
Record every suppression with its citation, so the trail exists before anyone asks for it.
Expire beliefs. "This is normal" is a claim about a system that changes. A baseline that has not been reconfirmed in ninety days should decay in confidence, not sit at full strength forever. Both mem0 and Hindsight describe temporal handling in retrieval; whether your decay policy is correct is a question their documentation cannot answer.
The third point is the one most likely to bite. An ADD-only store that never forgets will happily carry a stale baseline into a changed world, and retrieval-time temporal reasoning only helps if the time information is captured accurately on write.
Use case B — Churn prediction
A churn agent's memory problem is the opposite: choosing with a record. The volume is lower, the horizon is longer, and the outcome arrives weeks after the action. The failure mode is not noise but repetition — recommending a play that already failed for this account, or for accounts like it.
What to store
| Memory | Content | Layer | Written by |
|---|---|---|---|
| Trajectory | How the account moved, not just where it stands | Semantic time series | Extraction from usage and billing events |
| Interventions | Every play attempted, by whom, when | Episodic | CRM and outreach systems |
| Outcomes | What the play produced, scored once the horizon closed | Consolidated belief | Delayed scoring job |
| Segment effectiveness | Which plays work for accounts that look like this one | Procedural | Aggregation over scored outcomes |
| Relationship graph | Who left, who joined, which champion went quiet | Semantic | Entity extraction and linking |
The relationship graph is where cognee's ontology support and Hindsight's entity normalisation matter. "The admin left" is a fact about a person, an account, and a date simultaneously, and it only becomes predictive when those three are linked. A pure vector store flattens it into a sentence that happens to be retrievable.
The delayed-label problem
This is the hard part and it is not a memory-system feature, it is a design obligation. An intervention taken today has no outcome for sixty or ninety days. Until then:
Store the intervention as an episodic fact, not a belief. "We offered a discount on 3 March" is true. "The discount worked" is not yet knowable.
Promote to consolidated belief only after the horizon closes and the outcome is scored. This is where an ADD-only model helps: the intervention record and the later outcome record coexist, and retrieval with temporal reasoning assembles the pair rather than one silently overwriting the other.
Treat segment effectiveness as procedural memory with a sample size attached. "Discount play, 14 of 22 retained, EMEA mid-market" is actionable. "Discount plays work well" is folklore with a confident voice.
The flow
account crosses risk threshold
→ read standing answer: account trajectory summary # stored read
→ recall(interventions for this account) # bounded
→ recall(segment effectiveness for this profile) # bounded
→ recommend, citing prior attempts and the segment record
→ retain the recommendation as an episodic fact
→ [60-90 days] score the outcome, consolidate into belief
Every recommendation should surface what was already tried and what the segment record is. An agent that says "offer a discount" is guessing. An agent that says "offer a discount — this account has not had one, and the play retained 14 of 22 comparable accounts" is reasoning, and a human can disagree with it on the evidence.
Governance, privacy, and the things that end projects
Churn memory contains customer data. Anomaly memory contains infrastructure topology. Both are attractive to an attacker and both are in scope for a regulator.
Isolate at the store level, not the query level. Hindsight's banks are documented as isolated with no cross-bank leakage; cognee has a permissions model over datasets; mem0 scopes by user, session, and agent. Filtering by user_id at query time is a correctness measure, not a security boundary — a bug in filter construction becomes a data breach. Where the platform offers a store-level boundary, use it.
Scan on write, not on read. Hindsight's Memory Defense checks every retain against 45 secret and PII patterns and either redacts or blocks. Write-time scanning means the sensitive value never lands in storage. Read-time masking means it did land, and you are now relying on every future read path being correct. If your chosen system lacks this, build it into the write path yourself before the first production write, not after.
Memory is a prompt injection surface. This is the risk specific to agent memory and it is under-discussed. Anything the agent retains can be retrieved later and placed into a prompt. A hostile string in a support ticket that reaches churn memory becomes an instruction the agent reads with authority, weeks later, in a different session, with no obvious link back to its origin. Treat retained content as untrusted input at recall time, keep the provenance of every memory item, and never let retained text reach a position where it can be interpreted as system instruction.
Keep a deletion path. cognee exposes forget for an item or dataset. Whatever system you pick, know the answer to "a customer has exercised their right to erasure — what exactly do we run?" before you store the first customer record, because retrofitting deletion into a consolidated-belief store where the fact has been absorbed into a derived belief is considerably harder than it sounds.
Local-only is a real option. cognee runs text ingestion, retrieval, and session storage with no LLM key, using local GLiNER extraction and local embeddings. For an initial evaluation on sensitive data, that removes an entire category of approval conversation. Note the project's own framing: the bundled GLiNER extractor is a demo of its small-model pipeline, with a production-grade version available on request.
Evaluating memory, and what "better" means
Do not evaluate memory on public benchmark scores. Evaluate it on your workload, for four reasons: the benchmarks are conversational and your workloads are not, the published numbers carry vendor-specific configuration caveats in all three cases, the entity vocabulary in your domain is nothing like the benchmark's, and the thing you care about — did the agent stop repeating itself — is not what those benchmarks measure.
Build a small evaluation set from real history. Fifty anomaly investigations with known outcomes. Fifty accounts with known churn results and known intervention history. Then measure:
| Metric | Why it matters |
|---|---|
| Tokens per resolved task | The only figure connecting spend to outcome |
| Recall budget utilisation | If you never hit the ceiling, the ceiling is too high |
| Suppression precision | Of alerts suppressed, how many were genuinely benign |
| Suppression recall | Of genuine incidents, how many were wrongly suppressed — weight this heavily |
| Repeat-play rate | How often the churn agent recommends something already tried and failed |
| Citation validity | Sampled: does the cited evidence actually support the claim |
| Answer quality at each budget | Trimming is free until it silently is not |
Run the budget sweep deliberately. Try 2K, 4K, 8K, and 16K recall budgets against the same evaluation set. The curve almost always flattens well before the context window does, and the point where it flattens is your budget. Setting it by intuition wastes money in one direction and quality in the other.
All three projects support this properly: mem0 has open-sourced mem0ai/memory-benchmarks, Hindsight exposes Prometheus metrics for LLM calls, tokens, and latency, and cognee publishes its BEAM methodology with the reproduction gaps stated.
Best practices
| # | Practice | Why | Applies most to |
|---|---|---|---|
| 1 | Measure tokens per resolved task before adding memory | Without a baseline, every savings claim is unfalsifiable | Both |
| 2 | Set the recall budget as a hard ceiling and tune it with a sweep | The useful budget almost always flattens well before the context window | Both |
| 3 | Retrieve on several signals and fuse by rank | Identifiers such as hostnames and account IDs are lexical and entity signals, not semantic ones | Both |
| 4 | Store every belief with its evidence, proof count, and timestamp | An unauditable belief cannot be safely acted on | Both |
| 5 | Precompute standing answers for questions asked every session | It removes the work instead of making it cheaper | Both |
| 6 | Gate distillation on human or rule-based acceptance | One bad run should not permanently change the store | Both |
| 7 | Decay or reconfirm beliefs on a schedule | "Normal" is a claim about a system that keeps changing | Anomaly detection |
| 8 | Suppress only above a deliberate proof-count threshold, and cite the observation | A wrong suppression causes harm; a noisy alert only causes annoyance | Anomaly detection |
| 9 | Keep interventions as episodic facts until the outcome horizon closes | Otherwise the system learns from outcomes that have not happened yet | Churn prediction |
| 10 | Attach sample sizes to segment-level play effectiveness | "14 of 22 retained" is evidence; "works well" is not | Churn prediction |
| 11 | Isolate at the store level and scan for secrets and PII on write | A query filter is a correctness check, not a security boundary | Both, critical for churn |
| 12 | Treat recalled text as untrusted input and keep its provenance | Memory is a delayed prompt-injection surface | Both |
| 13 | Define the erasure procedure before the first customer write | Removing a fact after it has been folded into derived beliefs is hard | Churn prediction |
| 14 | Name an owner for the memory store | Unowned memory fills up with stale and orphaned data | Both |
Pitfalls
Implementing only episodic memory. Adds a retrieval bill without removing a reasoning bill. Costs go up and nobody can explain why. Consolidation and standing answers are where the return is.
No token budget. "Retrieve the relevant memories" without a cap reproduces the stateless cost curve with extra steps.
Vector-only retrieval. Hostnames, account IDs, and error codes are lexical and entity signals. All three systems combine signals for a reason.
Storing conclusions without evidence. An unauditable belief cannot be safely acted on, and the first review will say so.
Beliefs that never expire. A baseline from a system that has since been re-architected is worse than no baseline, because it is confidently wrong.
Ungated distillation. Automatic promotion of session lessons lets one bad run permanently poison the store.
Query-time isolation only. A filter bug becomes a data breach. Use the store-level boundary the platform provides.
Treating retained text as trusted. Memory is a prompt injection surface with a long fuse and no obvious attribution path.
Quoting benchmark numbers internally without the caveats. mem0's reflect a managed platform; cognee's 10M result is explicitly exploratory. Re-run them on your data or do not cite them.
Promoting churn interventions to beliefs before the horizon closes. The system then learns from outcomes that have not happened yet, which is the most expensive kind of wrong.
Building on a demo-tier deployment path. cognee's single-Postgres graph store is currently a demo feature. Check the tier before designing the topology.
No deletion story. Discovered at the first erasure request, which is the worst possible moment.
A phased plan
Phase one — instrument before you build. Measure tokens per resolved task for both agents as they are today. Without this baseline, every later claim about savings is unfalsifiable. Expect this to be the phase people want to skip.
Phase two — bounded recall only. Add episodic memory with a hard token budget and multi-signal retrieval. Run the budget sweep. This alone captures the quadratic-to-linear shift and is where most of the cost saving lives.
Phase three — consolidation. Introduce beliefs with attached evidence and proof counts. For anomaly detection, this is where suppression becomes possible. Measure suppression precision and recall before anyone acts on a suppression.
Phase four — standing answers. Precompute the questions asked every session. Cheapest reads in the system; largest remaining saving.
Phase five — gated distillation and procedural memory. Only now, with measurement in place and beliefs auditable, let the agents write durable lessons. For churn, this is where segment effectiveness becomes a measured asset rather than folklore.
Resist compressing this. Phases three through five are only safe because phases one and two made the system observable.
FAQs
Which of the three should we pick? Pick on your binding constraint, not on benchmark scores. If you need to evaluate on sensitive data with nothing leaving the perimeter, cognee's keyless local path is the only one of the three that offers it. If consolidated beliefs, standing answers, and write-time PII scanning are what you need — and for these two use cases they largely are — Hindsight has the most complete story. If you want the smallest possible integration surface onto an existing agent, mem0 is the shortest path. cognee's COGX import from mem0, Letta, Zep, and Graphiti means the first choice is not irreversible.
Can we use more than one? Yes, and for two use cases with different shapes it may be right — but only if you keep one system per use case rather than layering both on one workload. Two memory systems over the same agent means two write paths, two consolidation policies, and no single answer to "why did it say that."
Does memory actually reduce cost, or just move it? It moves some and removes some. Extraction on write and reranking on read are real costs. The removal comes from the quadratic term: material introduced early stops being re-billed every turn. Measure end to end, including the memory system's own model calls, and account for prompt caching on the stable prefix. The direction holds; the magnitude is yours to establish.
How does this relate to RAG? RAG retrieves from a corpus someone else curated. Agent memory retrieves from a corpus the agent itself wrote, including its own conclusions. The retrieval machinery overlaps heavily — both use hybrid search and rank fusion — but the write path and the governance model are different, and the write path is where the difficulty is. Hindsight positions itself explicitly against both RAG and plain knowledge graphs on this basis.
Should we fine-tune instead? Fine-tuning changes what the model knows in general. It cannot hold that account 88213 was offered a discount on 3 March, and it cannot be updated when that fact changes at 3pm. Use fine-tuning for behaviour and format, memory for facts and history. The two solve different problems and are not substitutes.
What happens when we change the underlying model? This is an argument for memory rather than against it. Facts and beliefs in an external store survive a model swap; everything held only in a long context window does not. Re-run the evaluation set after the swap — retrieval quality and the useful budget can both shift.
How much memory is too much? When retrieval precision starts falling, or when consolidation can no longer keep the belief set coherent. Both are observable if you instrumented in phase one. The corpus growing is not itself a problem; the corpus growing faster than consolidation compresses it is.
Who owns the memory store? Someone has to, and it is the question most likely to go unanswered. Memory that nobody owns accumulates stale beliefs, ungated distillations, and orphaned data with no deletion path. Name an owner in phase one.
References
mem0 — github.com/mem0ai/mem0; Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory, arXiv:2504.19413; evaluation framework at github.com/mem0ai/memory-benchmarks
Hindsight — github.com/vectorize-io/hindsight; arXiv:2512.12818
cognee — github.com/topoteretes/cognee; Markovic, Obradovic, Hajdu, Pavlovic, Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning, arXiv:2505.24478; BEAM evaluation methodology in the project repository
Edge et al., From Local to Global: A Graph RAG Approach to Query-Focused Summarization, arXiv:2404.16130
Cormack, Clarke, Buettcher, Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods, SIGIR 2009
Closing
The two agents in this article failed for the same reason and will be fixed by the same machinery, configured differently. The anomaly agent needs to remember in order to stay quiet. The churn agent needs to remember in order to choose well. In both cases the memory that matters is not the transcript — it is the small set of conclusions the system has earned, with the evidence still attached and a timestamp saying when it was last true.
The token saving is real and worth having, but it is the second-order benefit. The first-order benefit is that the agent stops starting from zero. An agent that begins each session already knowing what normal looks like, what has been tried, and what did not work is doing a different job from one that re-reads a transcript and hopes.
Build the measurement first. Add bounded recall second. Everything after that is only safe because those two exist.
