OpenZync
Menu

Sign In
Back to blog
engineeringAugust 14, 2026OpenZync Team22 min read

Your Agent Aced the Benchmark. It Still Doesn't Remember.

Share:
On this page

The uncomfortable result

Your agent scored in the nineties on the recall benchmark. Near-saturated, the benchmark equivalent of finished — retrieval from a mountain of past text, working with the reliability of a student who actually did the reading. Then you hand the same agent a realistic multi-session job: remember what it learned last week, and use it to act today. The score stops predicting anything.

That is the finding MemoryArena — accepted at ICML 2026 — produced with a benchmark built for exactly this question. Its thesis is one sentence, and it should bother every team shipping a memory-enabled agent: “agents with near-saturated performance on existing long-context memory benchmarks like LoCoMo perform poorly in our agentic setting, exposing a gap in current evaluations for agents with memory.”

The benchmark-reality gap: a benchmark score bar shows 94 with LoCoMo near-saturated, but the same agent fails a real multi-session task — a red cross over the real task and a sliver of a bar showing 12 percent recall on a fresh session

The nineties on recall and a fail on the task are not a contradiction; they are two different skills being measured. The benchmark asks whether the agent can find the fact. The task asks whether the agent can use the fact — mid-task, under pressure, when the next action depends on it. The field built years of benchmarks for the first question and almost none for the second, and the 2026 benchmark season made the difference impossible to ignore: agents that near-saturate LoCoMo-style recall benchmarks perform poorly on realistic multi-session agentic tasks.

Put a face on it, because the abstraction hides how ordinary the failure is. A support agent that scores in the nineties on recall can find the fact that Acme moved to the enterprise plan in March, driven by compliance — the exact strings, buried three weeks deep in a transcript. The score says nothing about whether that same agent, handed a fresh ticket asking why the bill is higher than expected, recognizes that the plan change is the reason, reaches for the migration status, and explains the charge instead of quoting a stale pricing page. One is retrieval; the other is the job. The benchmark measures the first and is silent on the second, which is how the nineties and the failure coexist without contradiction.

This post maps what the benchmarks actually measure, why memorization and action are coupled in a way the benchmarks miss, where the research is already moving — from static retrieve-then-reason toward reconstruction over temporal structures, from context stuffing toward memory as a structured process — and what a team that wants to measure memory honestly should build for itself.

What benchmarks actually measure

LongMemEval, the ICLR 2025 benchmark from Di Wu and colleagues, is the cleanest statement of what the field measures. It tests five abilities: “information extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention” — across “500 meticulously curated questions embedded within freely scalable user-assistant chat histories.” Its headline finding is a warning shot: “commercial chat assistants and long-context LLMs showing a 30% accuracy drop on memorizing information across sustained interactions.”

Each ability is a specific failure the field has watched agents make in the wild. Information extraction: pull the customer’s stated preference out of a chat. Multi-session reasoning: connect what they said last week to what they ask today. Temporal reasoning: know the plan they are on now, not the plan they were on. Knowledge updates: notice that the fact changed and answer with the new one. Abstention: say nothing rather than guess. Read as a list, it is the syllabus for a memory course — and it is exactly the list this series has been circling since why context windows aren’t memory: context is what the model sees, memory is what the system commits, and the five abilities are the ways committing can go wrong.

Read the list again and notice what the abilities have in common. Every one of them is a recall test: can the agent find the fact, place it in time, notice that it changed, know when to stay silent. The evaluation is a lookup, not a job. The agent is asked questions about a transcript while doing nothing — no actions, no consequences, no decisions that hinge on what it remembers. Retrieval is tested in a vacuum, and a vacuum is a generous environment: nothing is happening, nothing is at stake, no next action depends on the answer. The agent is never under the one condition that defines real memory — the need to act on what it remembers.

Finding a fact in a mountain of past text is a real and genuinely hard skill. Retrieval is not solved, and nobody in this series is pretending it is: LoCoMo — the long-conversation recall benchmark this series has referenced since the beginning — still leaves room at the top. The team-memory post noted that the Governed Memory paper reports 74.8% overall accuracy on LoCoMo against an 87.9% human baseline — even a state-of-the-art memory system leaves a double-digit gap to a human who actually remembers, and the humans themselves miss more than one in ten.

Recall is not done. It is simply not sufficient. Memorization in isolation is the easy half of memory; the hard half is what no recall benchmark shows: the agent that remembered, and then acted on it. The benchmarks measure the first half with increasing precision, and the second half not at all.

The coupling problem: memorization and action

MemoryArena’s diagnosis is precise. Its paper opens with the thesis: “Existing evaluations of agents with memory typically assess memorization and action in isolation.” In the benchmark world, the two never meet: memorization is tested by a question set, action is tested by a task suite, and nothing in between asks whether memory guided the action that followed it.

The paper’s response is “a unified evaluation gym for benchmarking agent memory in multi-session Memory-Agent-Environment loops” — the loop being the structure real agents actually live in. The agent acts, the environment responds, the agent distills what it learned into memory, and that memory guides the next action. Break any link in the loop and the loop dies; the benchmark is built so the whole loop has to work.

The loop is not a metaphor; it is the runtime. Take bundled web shopping: the agent shops across sessions for a list of items under learned preferences, and each purchase is a write to memory that must steer the next one. Preference-constrained group travel planning: the agent learns in session one that one traveler refuses red-eye flights, and in session three that constraint must steer the route proposal without being re-stated. Progressive information searching: the agent accumulates leads across sessions and prunes them as new evidence arrives. Sequential formal reasoning over math and physics problems: the conclusion of step one is the premise of step two, and forgetting it is not a miss — it is a wrong answer.

The gym spans four domains: “bundled web shopping, preference-constrained group travel planning, progressive information searching, and sequential formal reasoning over math and physical problems.” The tasks are “human-crafted agentic tasks with explicitly interdependent subtasks, where agents must learn from earlier actions and feedback by distilling experiences into memory, and subsequently use that memory to guide later actions to solve the overall task.” (arXiv:2602.16313, He et al.; the paper was also presented at the KnowFM workshop at ACL 2026.)

The word that matters is interdependent. The subtasks are wired together: what the agent does in step two depends on what it learned in step one, which means the answer cannot be pre-loaded, cannot be retrieved from a transcript, cannot be found with a lookup. The memory has to be written mid-task — the agent must recognize what is worth keeping in the noise of a working session — and then used mid-task, at the moment the next action depends on it. Near-saturated recall says nothing about either skill. A near-saturated recall score says the agent can find a fact that is already in the store; it does not say the agent can build the store in the first place, or reach for the right memory at the right moment, or know which of the facts it holds is the one that applies now.

The failure mode is concrete. A recall-saturated agent, dropped into the loop, does what it was trained to do: it waits for a query, retrieves a snippet, answers. But in the loop there is no query waiting for it — there is an environment responding to its actions. It must decide what to write before it can retrieve anything, and it must decide what to act on before it has been asked anything. The benchmark never asks that question, which is why its scores do not survive contact with the loop.

Why recall is not memory

The deeper problem is the paradigm. MRAgent, a graph-memory paper accepted at ICML 2026 (Ji, Li, and Hooi), states the critique directly: “current memory-augmented agents rely on a static retrieve-then-reason paradigm, this rigid pipeline design prevents them from dynamically adapting memory access to intermediate evidence discovered during inference.”

Static retrieve-then-reason: query the store once, take the top-k, reason over whatever came back. The pipeline is fixed before the agent starts working. But real tasks are not fixed — the agent discovers things mid-task that should change what it retrieves — and a pipeline that cannot adapt to intermediate evidence is a pipeline that answers with yesterday’s view of the situation. The rigidity is not a performance issue; it is a structural one. The pipeline has no place for the loop to feed back into, so whatever the agent learns while working cannot change what it looks at next. The memory and the action run on different tracks, and the benchmark scores the track that never touches the work.

Retrieving the right snippet is not the same as reconstructing a situation-specific view. A snippet is a fragment of the past; a view is the past assembled for the present — what was true then, what is true now, what changed in between. And the difference is not cosmetic, because a wrong-but-retrievable fact is worse than no fact at all. A fact that is retrieved is treated as ground truth by the agent that retrieved it — which is precisely why this series has spent two posts on what happens when the store is wrong. The poisoning post showed what a stored wrong fact does: it does not fail to help, it actively steers every future interaction that retrieves it — poison once, exploit forever. The team-memory post showed the governance half: the correction loop that would fix a wrong fact is the one primitive nobody ships. A benchmark that only measures whether the fact was found cannot see any of this — a confidently wrong fact scores exactly the same as a correct one, as long as both are retrieved.

That last point deserves to be said plainly, because it is the quietest failure in the whole picture. Retrieval quality is measured against the store, not against reality: the benchmark rewards the agent that faithfully returns a stale fact exactly as much as the agent that returns a current one, because both were “found.” The agent cannot be penalized for acting on a wrong fact the benchmark never asked it to act on. The score is a measure of the store’s reach, not of the agent’s judgment — and judgment is where the memory-action loop actually lives.

Recall is a measure of the store. Memory is the whole loop: what gets written, what gets kept, what gets corrected, what gets used. The benchmarks measure the store and call it memory, and the gap between the two is the gap in this post’s title.

The efficiency wave: memory as a structured process

Two ICML 2026 posters, from opposite ends of the stack, converge on the same lesson: memory is not context stuffing. Both treat memory as a structured process, and both measure the payoff in scores and in tokens.

SimpleMem, from Jiaqi Liu and colleagues, is a three-stage pipeline: “(1) Semantic Structured Compression, which distills unstructured interactions into compact, multi-view indexed memory units; (2) Online Semantic Synthesis, an intra-session process that instantly integrates related context into unified abstract representations to eliminate redundancy; and (3) Intent-Aware Retrieval Planning, which infers search intent to dynamically determine retrieval scope and construct precise context efficiently.” The results are the point: “achieving an average F1 improvement of 26.4% in LoCoMo while reducing inference-time token consumption by up to 30×.” (code)

Read the three stages as three answers to three questions the retrieval-only view never asks. What is worth keeping? — compression decides, distilling unstructured interactions into indexed units instead of storing them whole. What is redundant? — synthesis decides, merging related context in-session so the store does not accumulate the same fact five times. What should the agent look at? — retrieval planning decides, inferring intent to scope the search instead of throwing the whole store at the prompt. Each stage is a structure the static pipeline lacks, and each one independently moves the score.

EAM — Executable Agentic Memory, from Zerui Qin and colleagues, applies the same instinct to GUI agents. Instead of free-form generation over screenshots, it builds “a structured Knowledge Graph (KG) that shifts GUI planning from free-form generation to a robust retrieval-and-execution process,” driven by “a value-guided graph search where a lightweight Q-function model steers Monte Carlo Tree Search (MCTS) over the KG.” The numbers again: “EAM beats UI-TARS-7B by up to 19.6% on AndroidWorld, while reducing token costs 6× relative to GPT-4o. With a 2.8s average latency.”

One pipeline compresses conversation into indexed units and plans retrieval by search intent; the other compresses GUI state into a knowledge graph and plans actions by search over it. Different stacks, identical direction: memory is compression, structure, and intent — not a bucket to dump context into. And in both cases the structure pays twice: better scores, and an order-of-magnitude cut in inference cost.

That economics matters more than it looks. A memory system you cannot afford to run is a memory system you will not run — and the 30× and 6× figures are not optimizations, they are the difference between memory as a line item and memory as an architectural assumption. When the cost of remembering collapses, the design question stops being “how much can we afford to store” and becomes “what is worth storing at all” — which is the question the benchmarks are finally starting to ask.

Reconstruction over retrieval

MRAgent names the paradigm shift in its title: Memory is Reconstructed, Not Retrieved. The framework: “we propose MRAgent, a framework that combines an associative memory graph with an active reconstruction mechanism.” The representation: “We represent memory as a Cue-Tag-Content graph, where associative tags serve as semantic bridges connecting fine-grained cues to memory contents.” The mechanism is the interesting part — reconstruction replaces the one-shot lookup with an iterative explore-and-prune loop over the graph, adapting access to what the agent discovers mid-inference, exactly the adaptation the static retrieve-then-reason paradigm rules out. The results: “Experiments on the LoCoMo benchmark and LongMemEval benchmark demonstrate significant improvements over strong baselines (up to 23%), while substantially reducing token and runtime cost.” (arXiv:2606.06036)

Two memory paradigms: retrieve is a one-shot static pipeline — query, vector store, static snippets, answer — while reconstruct adds an iterative explore-and-prune loop over a temporal graph, building a situation-specific view before answering

Retrieval returns fragments; reconstruction returns the situation. That is the whole difference, and it is why the memory-action gap is not an implementation detail — it is the difference between an agent that quotes the past and an agent that reasons from it. The top lane of the diagram is what the benchmarks measure: query, fetch, answer. The bottom lane is what agents actually do when the answer is not pre-loadable: query, explore, prune, build a view, then answer. Notice that the bottom lane is a loop — the graph feeds the reconstruction, and the reconstruction feeds back into the graph — which is the shape the benchmarks’ transcript-and-question format never exercises.

This is where the temporal-knowledge-graph architecture points — the design this series covered in how a temporal knowledge graph is built. A graph with temporal edges — facts carrying validity windows instead of overwriting each other — is exactly the structure reconstruction needs. It lets an agent reconstruct “what was true then, what is true now” instead of fetching static snippets that cannot distinguish a fact from the fact it replaced. Vector top-k resurface the stale fact alongside the current one; a temporal graph closes the stale window and returns the current fact, with the history still queryable. The question “what does this situation look like right now” becomes answerable by construction, not by luck:

Why the same query returns different answers: vector top-k resurface stale facts, while a temporal knowledge graph returns only the currently-true fact and keeps the closed window queryable as history

OpenZync is one architecture working in this direction: a self-hostable temporal knowledge graph where facts supersede rather than overwrite — reconstruction over retrieval, in the MRAgent direction, made operational. It is a starting point, not a claim of completion; the governance half of the same design is the subject of the team-memory post.

MemEvolve and the meta-frontier

Then there is the layer above the memory system: the memory system itself. MemEvolve, from Guibin Zhang and colleagues, is “a meta-evolutionary framework that jointly evolves agents’ experiential knowledge and their memory architecture, allowing agent systems not only to accumulate experience but also to progressively refine how they learn from it.” Not just better memory — memory that improves its own way of learning, “improving frameworks such as SmolAgent and Flash-Searcher by up to 17.06%,” with “strong cross-task and cross-LLM generalization.”

Underneath it sits EvolveLab, “a unified memory codebase that distills twelve representative memory systems into a modular design space (encode, store, retrieve, manage)” — twelve memory systems, reduced to four operations. That modularization is the tell: the field now has enough memory architectures to taxonomize them, which is what a field does right before its architectures start improving each other. The four operations are the same stages every memory system in this post already has — encode, store, retrieve, manage — and the difference is that MemEvolve evolves how those stages are wired, per agent, over time. The memory system is no longer a fixed design chosen at build time; it is a configuration the system itself refines as it accumulates experience.

The frontier is not better memory. It is memory that adapts its own structure — a system that learns not only the facts, but how it stores, retrieves, and prunes them. That carries an uncomfortable corollary for evaluation: if memory architecture itself becomes a moving target, then a benchmark that assumes a fixed store is measuring a system that no longer exists. An eval suite built around a static retrieve-then-reason baseline cannot score an agent whose memory reconfigures itself between runs — which is one more reason the eval gap is not going to close by improving retrieval alone. The evals have to evolve with the architectures — which is the subject of the table below.

The benchmark table

The 2026 season’s benchmarks, in one table. Read it as a map of what each evaluation can see — and what it is blind to:

BenchmarkWhat it measuresWhat it misses
LoCoMoLong-conversation recall — finding facts in a long transcriptAction coupling: nothing the agent does after remembering
LongMemEvalFive recall abilities: extraction, multi-session reasoning, temporal reasoning, updates, abstentionSingle-session recall focus: a transcript to look up, not a loop to act in
MemoryArenaMulti-session memory-action loops over interdependent subtasksEarly stage: four domains, one gym, results still settling
AndroidWorldGUI task completion on a phone emulatorDomain-specific: mobile GUIs, not general agent memory
OpenZyncTemporal knowledge graph with supersede-not-overwriteNo benchmark scores claimed; reconstruction over retrieval

Read the table as a map of the category, not a leaderboard: every benchmark measures a different slice, and the action slice — what the agent does with what it remembers — is nearly empty. LoCoMo and LongMemEval are recall tests over transcripts; neither has an agent acting. AndroidWorld has an agent acting, but only on a phone screen — its lessons about memory are real and narrow, and they do not generalize to an agent that reads tickets, remembers customers, and reasons across weeks. MemoryArena is the first serious attempt to put both halves together, and it is early — four domains, one gym, and a headline finding that is really an indictment of everything that came before it. The empty slice is exactly where the gap lives: the nineties on recall, the fail on the task. Every row in the “misses” column is a requirement the next eval harness will have to carry.

What to watch

Four signals will mark the category’s maturity. None of them are product announcements; all of them are evidence.

Multi-session eval harnesses become the standard. MemoryArena is the early shape of what evaluation looks like when it grows up: a gym, not a question set. The tell to watch is adoption — when multi-session memory-action loops appear in the eval suites teams run before shipping, the gap stops being a research finding and becomes a release gate. The paper is honest about its scope: four domains, human-crafted tasks, early days. The shape is the story, and the shape is a loop. The moment a benchmark that looks like a gym — with sessions, an environment, and interdependent subtasks — shows up in an enterprise eval suite, the category has turned a corner.

Forgetting becomes a first-class capability. Over-retention is a failure mode. A stale fact misleads more than an absent one: an agent that does not remember is visibly uncertain, while an agent that remembers yesterday’s truth is confidently wrong — and a confidently wrong answer, as this series has argued since the poisoning post, is indistinguishable from a correct one to everyone who reads it. LongMemEval already measures abstention as one of its five abilities — knowing when not to answer. The next step is forgetting-when-beneficial: knowing when to stop trusting a fact, retiring it before it misleads. The temporal rollback theme of the poisoning post — retire, don’t delete — is exactly the mechanism this capability needs: a fact can only be retired if the store kept the timeline that says when it stopped being true.

Cost-aware metrics become part of the eval. A 30× token cut and a 6× token cut do not just make memory cheaper; they make memory economics a benchmark dimension. A system that scores two points lower and costs thirty times less is a different engineering trade, and the field has started pricing it. Watch for eval harnesses that report tokens per correct answer, not just accuracy — the efficiency wave has already made the metric inevitable, because a score with no cost attached is a score nobody can act on.

Benchmark diversity beyond recall. One recall benchmark is not a category. The 2026 season spans transcript recall (LoCoMo, LongMemEval), action coupling (MemoryArena), and GUI completion (AndroidWorld); the maturity signal is when a single agent’s eval suite deliberately mixes them, because production agents do all three at once. A support agent that reads a ticket, remembers the customer, and acts on a GUI is all three benchmarks at once — and no single one of them measures it. A harness that mixes the three is the first honest approximation of the job.

The eval-harness checklist, short version, for teams that want to measure the gap before it measures them:

  • Test multi-session dependencies: does a fact learned in session one change what the agent does in session five?
  • Test hidden dependencies: tasks where the memory the agent needs is discoverable only mid-task, not pre-loadable.
  • Test stale-fact handling: does the agent notice when the fact it retrieved stopped being true?
  • Measure token cost: what each correct answer actually costs, not just how often it is correct.
  • Measure forgetting: whether the agent drops what it should — and knows it did.

Five checks, none of them exotic, all of them invisible to a recall benchmark. The last one is the least tested and the most telling: forgetting is the only check that requires the store to keep a timeline, which is the same requirement every other honest memory evaluation keeps circling back to.

The eval harness is the next feature

Benchmarks measure memorization in isolation. Agents work in memory-action loops: they acquire memory while interacting, then use it to guide future actions. Those are different skills, and the 2026 benchmark season made the difference measurable in one sentence — near-saturated recall, poor agentic performance, from a benchmark that finally put the two together. The research is already moving: reconstruction over retrieval (MRAgent), memory as a structured process (SimpleMem, EAM), memory that evolves its own architecture (MemEvolve), and evaluation that finally includes the loop (MemoryArena’s gym). What is missing is the habit: teams measuring memory the way they measure latency — with a harness that tests dependencies, staleness, cost, and forgetting, before shipping.

If you are building a memory-enabled agent, that checklist is the start, and the storage design underneath it matters more than the checklist’s position in the file tree. A store that overwrites in place cannot answer “what was true then,” cannot retire a stale fact, cannot audit what the agent believed when it acted — which means it cannot be evaluated on three of the five checks at all. OpenZync is a self-hostable starting point for the storage side: a temporal knowledge graph where facts supersede rather than overwrite, exposed over REST and MCP — memory built so the timeline lets you reconstruct what was true then and what is true now, which is what makes it evaluable in the first place. Start with the memory & context docs and the openzync-mcp repository.

This is part of a series on agent memory. Read why context windows aren’t memory, then how a temporal knowledge graph is built, then the MCP memory-server field guide, then when agents talk to each other, who remembers, then the new attacks on AI memory, then team memory and who fixes a wrong fact.