The uncomfortable result
Your agent does not need a knowledge graph. It needs a search box.
That sentence sounds like provocation. As of this month, it is a benchmark result: a preprint released in August 2026 handed an agent nothing but a lexical index over its raw conversation archive and an iterative search loop — no summaries, no embeddings, no graph — and the agent answered precise-retrieval questions more accurately than the graph- and tree-based memory systems it was compared against. The result is uncomfortable for anyone who picked a memory architecture this year, and the uncomfortable part is not that the search box won. It is how much of the structured systems’ advantage was never retrieval at all.
Put a face on it, because the abstraction hides how ordinary the surprise is. A support agent with a memory stack built this year was given a knowledge graph over every conversation it has ever had — entities extracted, relations drawn, communities clustered — and a second agent was given the same conversations untouched, with a search box it controls itself. Ask both “what plan is this customer on?” and the search box answers first, and answers right. The team that paid for the graph does not have a better agent; it has a slower one that cost more to build. That is the uncomfortable part: the structure did not fail, it just never pulled its weight at the one job it was hired for.
The instinct is to explain the result away — one paper, one benchmark, one interface built by its own authors. The instinct is worth resisting, because the paper is careful in exactly the places a hype cycle is not: it matches backbones, it reuses baselines instead of inventing strawmen, and it says plainly what it did and did not measure. The question it leaves on the table is the one this post keeps returning to: if retrieval over the raw record is this cheap, what was the structure for?
The paper is “When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory” (Li et al., arXiv:2608.12888), a preprint from August 2026. Its system, ReFind, builds no semantic structure at all, and it out-retrieved the graph- and tree-based memory systems in the comparison on the exact kind of question this series has cared about since why context windows aren’t memory: precise, evidence-grounded questions over chat archives.
This post walks what that result does and does not say, what structure actually costs, and why the honest conclusion is not “search beats graphs.” It is that retrieval and truth-resolution are two different jobs, and structure only earns its keep on the second one. The synthesis is a two-layer architecture: keep the raw record as the retrieval layer, and put a temporal layer on top for the questions a search box cannot answer.
The default assumption
Every serious agent-memory project this year started from the same assumption, and ReFind states it exactly. The paper’s opening sentence is its thesis: “Agent-memory systems increasingly buy retrieval quality with structure, transforming raw conversation histories into summaries, embeddings, trees, or knowledge graphs before any question is asked. We ask how much of that benefit comes from the structure itself, rather than from competent retrieval over the raw history.”
That thesis is the default bet of LLM memory systems — and the default assumption beneath most agent memory architecture decisions in 2026: structure is what makes memory retrievable, so the field pays a construction bill in exchange for retrieval quality. Build the summaries before the questions arrive. Embed everything. Materialize the graph. The structure is the product; the raw history is the feedstock. The assumption felt safe because it is not one bet but a cascade of them — extraction quality, entity resolution, edge construction, retrieval over the built artifact — and each stage looked like progress in isolation. Nobody questioned the first domino: that any of it was necessary to find a fact in a conversation.
The structure-building wave is not hypothetical; it is the 2026 literature. MRAgent, accepted at ICML 2026, builds a Cue-Tag-Content graph over memory, with associative tags serving as semantic bridges between cues and contents. SimpleMem, also ICML 2026, compresses interactions into compact, multi-view indexed memory units and plans retrieval by search intent. EAM, ICML 2026 as well, builds a knowledge graph over GUI state and steers action planning over it. Three different stacks, one shared premise: transform the history before the question, because that is where retrieval quality is assumed to live. This series has covered all three as part of the efficiency wave — they cut token costs and move scores — but note what they all have in common: not one of them asks whether the raw history, searched competently, would have moved the same scores.
That premise is the common ancestor of the flood of “vector database vs knowledge graph” comparison guides that 2026 produced — every one of them asks which structure to build, none of them asks whether to build one at all. It is also the premise behind this series’ own product, in how a temporal knowledge graph is built. ReFind is the first paper this series has covered that holds the premise up to measurement — and it deserves a serious reading precisely because it shares the wave’s assumptions about everything except structure itself.
The search box that won
ReFind is a deliberate refusal to build. Its description is the design: the system “leaves the conversation archive unmodified, indexes it lexically at turn granularity, and combines a generic iterative keyword-search loop with four chat-native controls grounded in empirical refinding work: session-aware rank fusion, local context expansion, temporal narrowing, and skipping already-inspected sessions.” No LLM-based index construction at all — no summarization pass, no embedding projection, no graph build. The index is lexical. The search is keyword search, iterated by the agent rather than answered in one shot. And the four controls are the things a human actually does when refinding a conversation: merge the best hits across sessions, expand the context around a match, narrow by time, and skip the sessions already inspected. That is the entire system — roughly one evening of engineering and one morning of indexing, versus the weeks of pipeline work a structured memory stack demands.
What the paper calls agent-controlled search is the difference between a lookup and a loop. A retrieval stage answers once: query, top-k, done — the static retrieve-then-reason pipeline this series traced through the gap between benchmark scores and agentic memory. ReFind answers iteratively: search, read what came back, narrow by time, skip what was already inspected, expand around the promising hit, search again. The agent is in the loop, deciding what to look at next, and the four chat-native controls are the vocabulary of that decision. The difference sounds modest and is not: the system adapts its access to the evidence as it reads, which is precisely what a one-shot top-k cannot do.
The comparison is uncompromising: “roughly 2,800 questions on precise-retrieval and fact-tracking capabilities evaluated under the incremental multi-turn setting of MemoryAgentBench, ReFind attains the highest mean accuracy (58.2) of any system compared, above the strongest graph- and tree-based memory systems (HippoRAG 2, 53.2), all under a GPT-4o-mini backbone matched to every reused baseline.” The matched backbone matters: the baseline systems were not handicapped, and the search box was not given a stronger model to compensate — the same reasoning model, the same retrieval stage replaced. On the long-conversation benchmarks the gap widens further: “On LongMemEval-S/M, the same interface reaches 93.2 ± 3.3 and 89.3 ± 6.0 with GPT-5-mini.”
The punchline is the paper’s own verdict:
“The results indicate that for precise, evidence-grounded questions over chat archives, much of the benefit credited to elaborate memory structures is recoverable by giving an agent controllable search over the unmodified record, with no LLM-based index construction at all.”
Two phrases in that verdict deserve emphasis. “Much of the benefit” — not all of it, and the paper says so on purpose. And “for precise, evidence-grounded questions over chat archives” — a scope, not a universe. Both limits are the subject of the next section, because the distance between the honest reading and the marketing reading of this paper is where the architecture debate actually lives.
What the result does and doesn’t say
Read the result as narrowly as the paper does. The claim is scoped to precise, evidence-grounded questions over chat archives — the retrieval stage of answering, not answering itself. ReFind collects the evidence, and “a separate reasoning stage answers from the collected evidence,” so the benchmark isolates the store and the search from the model doing the reading. Every system in the comparison runs under “a GPT-4o-mini backbone matched to every reused baseline,” so no baseline was handicapped by a weaker model. The tasks are “single- and multi-hop QA, event ordering, and fact consolidation,” evaluated under the incremental multi-turn setting of MemoryAgentBench — not open-ended conversation, not action, not judgment. That is the boundary of the claim, and the paper does not pretend otherwise.
Two over-reads are tempting, and both are wrong. The first is “graphs are dead.” Nothing here says that. What died is the assumption that structure is required for retrieval: the graph and tree systems were beaten at their own game by a lexical index and a search loop, but the game is one slice of memory — finding a fact in the record. The graph’s other jobs — resolving what was true then versus now, preserving who said what with what standing, reconstructing what the agent believed when it acted — are untouched by this benchmark, because they are not retrieval jobs at all. The second over-read is “graphs never helped.” Also wrong. This benchmark says nothing about the questions structure was built for in the first place; a graph that fails at keyword retrieval may still be the only thing that answers a temporal question. ReFind did not test those, and no search box claims to.
Put the boundary to work. “Did the customer mention a migration in March?” is a precise, evidence-grounded question over a chat archive: search answers it, and answers it well. “Are they still on the plan we moved them to in March?” is the same words and a different job — the answer depends on what has happened since, on supersession, on a timeline the raw record does not rank by. The first question is ReFind’s benchmark; the second is the reason the temporal layer exists. Conflating the two is how teams buy structure for the first question and then wonder why the second still fails.
The honest middle is where the architecture debate should live. For retrieval over the raw record — the “knowledge graph vs RAG” question every guide asks — competent agent-controlled search closed most of the gap, at a fraction of the construction cost. That shifts the burden of proof onto the structured approach: its case now has to be made on the questions search cannot answer, which is the subject of section six. This is the same lesson as the gap between benchmark scores and agentic memory: a benchmark measures a slice, and the slice here is retrieval — precise, evidence-grounded, over the archive. The slice it does not touch is truth over time.
The cost of structure
Structure has a price, and a second preprint released this month priced it. “Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems” (Pollertlam and Kornsuwannawit, arXiv:2608.11879, August 2026) benchmarks “three memory systems (Mem0, Hindsight, and Mastra Observational Memory) against two reference strategies — a fixed-size rolling window and resubmitting the full transcript — across two backbones and conversations of up to 400 turns, pairing every cost measurement with answer accuracy on 665 LoCoMo questions.”
Three findings, each one a surprise.
Cost cannot be predicted from input size. The paper’s first finding is that “a memory system’s serving cost cannot be predicted from conversation length and message size alone: a regression that tracks the two reference strategies closely misses the memory systems by 18-69%, their cost driven instead by internal memory behavior.” You cannot budget a memory system by counting the tokens you feed it. The internal machinery — what it consolidates, when, how often — dominates the bill, and it is invisible from the outside until the invoice arrives. A system that looks identical on paper to another can cost half again as much to serve, and the difference is not in the conversation, it is in the structure.
The payback is not guaranteed. The paper finds that “whether — and when — a memory system becomes cheaper to serve than the full transcript is highly sensitive to the system and the backbone, from the first tens of turns for the cheapest to never within 400 turns for the most expensive.” Read that twice. Four hundred turns — a long conversation, a month of a busy support thread — and the most expensive system never breaks even against resubmitting the entire transcript, which costs nothing to construct and nothing to maintain. The cheapest systems pay back inside the first tens of turns, which is a real advantage — but it is an empirical property, not an architectural guarantee, and it flips depending on the backbone you chose.
Nobody wins both axes. The paper is blunt that “no system wins on both axes: accuracy spans 21-54%, and the backbone choice drives cost as much as the memory system does.” The accuracy spread across systems is wider than the gap between most of them, and the backbone moves the bill as much as the memory system does. Choose the wrong model and you pay the memory tax twice: once in tokens, once in answers you cannot trust.
Structure is not free. It carries a construction bill — the summarization pass, the graph build, the embedding projection — and a serving bill, and the break-even curves above are the honest accounting. The operational consequence is concrete: a team budgeting its memory layer by conversation volume is budgeting wrong, and a team picking a backbone without re-pricing the memory layer on top of it is choosing blind. Combined with ReFind, the picture is uncomfortable: the retrieval benefit structure pays for is largely recoverable without it, and sometimes the structure never pays back at all. The case for structure has to be made on other ground — which is the subject of the next section.
What search can’t answer
The search box wins on retrieval. The honest question is what retrieval cannot do — and the defense of structure lives entirely on the other side of that line. Four capabilities, each one a reason the raw record alone is not memory.
What was true then vs now. A log is a perfect record and a poor answerer. The archive preserves everything — including the fact you corrected three weeks ago — and search will happily resurface both versions, because both are in the record and both are equally findable. Current truth is not a property of the log; it is a property of a timeline that knows which fact closed when. That is supersede-not-overwrite — a new fact closes the old fact’s validity window instead of erasing it — and no amount of search over the raw record reconstructs it:
The support agent asks “what plan is the customer on?” and the search box returns both the enterprise upgrade and the plan it replaced, with equal confidence, because the log treats them as equally real. Search can rank by relevance; it cannot rank by recency of truth. The vector store has the same blind spot — top-k resurfacing serves the stale fact alongside the current one, and the reasoning stage cannot tell which is which without a timeline it was never given. Relevance and recency are different orderings, and only one of them is a function of truth.
Authority. Search returns the fact; it does not return the fact’s standing. Who said it, with what scope, under what conditions — the raw record carries the information, but nothing in keyword search weighs it. Worse, the consolidation layer actively strips it: the authority-collapse preprint found that consolidation “preserves a claim while erasing the source constraints governing its authorized use” (Zhan et al., arXiv:2608.01679, August 2026) — and that persisting authority labels cut the unauthorized-action rate from 16.9% to 0.0%. A fact distilled from a document marked “for internal review only” gets acted on as if it were public policy, because the claim survived the distillation and the permission did not. This is the deeper layer of the confabulation post: a memory that keeps the claim but loses the permission is a memory that manufactures authority.
Audit. When the agent acted on a wrong belief, what did it believe? A log can prove what the agent retrieved; it cannot show what the agent held to be true at the moment of action — the belief, not the retrieval. That distinction is the spine of the poisoning post: a stored wrong fact does not fail to help, it steers every future interaction that retrieves it, and the only way to find the wrongness is a timeline that records the belief as it changed. Compliance asks “what did the agent believe when it approved this?” and the raw record answers with a retrieval timestamp; the temporal layer answers with the belief itself, supersessions included.
Retirement. A fact does not stop being findable when it stops being true — it stops being true, and search cannot tell. A pricing fact superseded in April still tops the ranking in August, and the agent that cites it is not retrieving badly, it is retrieving accurately a fact that should not be trusted. Knowing when to stop trusting a fact is the forgetting capability both the team memory post and the gap between benchmark scores and agentic memory circled back to: retirement requires the timeline, and the timeline requires structure that supersedes rather than overwrites.
Four capabilities, one requirement each: a timeline, a source label, a belief history, a retirement mechanism. None of them are retrieval problems. All of them are structure problems — which is exactly the division of labor the next section proposes.
The synthesis: keep the log, add the layer
The architecture that survives both preprints is two layers, not one. The bottom layer is the record itself: the conversation archive, unmodified, indexed lexically at turn granularity, searchable by an agent that controls its own search loop — ReFind’s design, and the reason teams can stop building structure for retrieval. The top layer is the truth layer: a temporal index over the record that resolves current truth — what is true now, who said it, when it stopped being true — built from the supersession machinery this series described in how a temporal knowledge graph is built. Search retrieves; the temporal layer resolves. Neither substitutes for the other, and the two preprints define the boundary between them cleanly: ReFind shows that retrieval needs no structure; the four capabilities above show that truth-resolution needs nothing but.
The truth layer does not replace the log; it sits on top of it, and the log remains the retrieval substrate. An agent working against this stack searches the raw record first — fast, cheap, unmodified — and consults the truth layer only when the answer has to be current, authorized, or defensible. That is a smaller structure than the field has been building, doing a narrower job, and it is exactly the structure the preprints vindicate: the construction bill is paid only where search cannot reach. That is the direction OpenZync works in: a self-hostable temporal knowledge graph where facts supersede rather than overwrite, keeping the timeline a search box cannot reconstruct — the truth layer in the diagram, made operational. It is a starting point, not a claim of completion; where it sits among the other tools in this category is mapped honestly in the openzync-vs-zep-vs-mem0 post.
The decision table
The comparison guides keep asking which layer to build first. The honest answer is a table, because the choice is not “vector database vs knowledge graph” — it is which job each layer is being asked to do:
| Layer | Best at | Fails at | Build when |
|---|---|---|---|
| Raw logs + agent search | Retrieval precision over the unmodified record | Current-truth resolution; no timeline | Questions are evidence-grounded and the record is the asset |
| Vector database | Fuzzy similarity over unstructured text | Stale-fact resurfacing; no timeline | Semantic search at scale over text that resists structure |
| Knowledge graph | Relationships and structure | Raw retrieval when search suffices | Entity-dense, relationship-heavy domains |
| Temporal knowledge graph | What-was-true-then and supersession | Raw retrieval when search suffices | Agents whose past beliefs matter |
| OpenZync | Temporal knowledge graph with supersede-not-overwrite | Raw retrieval when search suffices | Provenance and audit; no benchmark scores claimed |
Read the rows as jobs, not as products. The first row is ReFind’s entire case: for finding facts in the record, the unmodified log plus a search loop is the retrieval layer, and every other row concedes that job to it in its own way — the vector store fails on stale-fact resurfacing, the graph fails on raw retrieval when search suffices, and the temporal graph’s claim is truth, not retrieval. The last row is this series’ own bet: the temporal layer is not a retrieval upgrade; it is the layer that answers what-was-true-then, keeps provenance, and makes belief auditable — and it claims no benchmark scores for retrieval, because the preprints above just redrew that map. Choose the row by the question the agent actually has to answer, and the guide pages start answering themselves.
What to watch
Four signals will mark where the category lands. None of them are product announcements; all of them are evidence.
Retrieval and truth-resolution split into separate benchmark axes. The 2026 eval season is quietly separating two skills that used to share one score. ReFind-style evals measure recall over the record — how well an agent finds what is there; supersession and abstention evals measure truth — whether the agent knows what is true now and when to stay silent. The maturity signal is when the two stop sharing a headline number, because the preprints above show the architectures differ by exactly that split: a search box wins the first axis and cannot touch the second; a temporal graph is the reverse. The gap between benchmark scores and agentic memory mapped the first crack in this wall; these two preprints are the wall coming down.
Cost-aware evaluation becomes the standard. “Total Recall at What Cost?” is the early shape of what evaluation looks like when it prices memory: cost and accuracy measured on the same questions, in the same run, not reported in different documents. The tell is the metric from that post’s checklist — tokens per correct answer — becoming the headline number. A score without a serving cost is a score nobody can budget against, and the 400-turn break-even curve is the reason: once a benchmark reports the price of its own answers, teams stop choosing architectures by accuracy alone. The two preprints above are already that change in miniature; the signal is when it stops being remarkable enough to blog about.
Search-native memory UIs emerge as a category. The four controls ReFind names — session-aware rank fusion, local context expansion, temporal narrowing, skipping already-inspected sessions — are not implementation details; they are the product surface of memory as refinding. Teams that give agents control over their own search are building a category, and the tell is when search over the raw record ships as a capability users see — an agent that visibly narrows by time, skips inspected sessions, expands around a match — not a pipeline they are told to trust. The refinding vocabulary is the product vocabulary.
Structure is justified by evidence, not convention. The burden of proof has flipped. A team proposing a graph or a summary pipeline now has to answer the question ReFind made unavoidable: what does this structure buy that search over the record cannot? Supersession, authority, audit, retirement — the four answers in section six — are the ones that survive. Structure built for retrieval is now a convention in search of a justification, and the preprints just made that convention expensive. The category matures when structure proposals come with a measured answer to that question, and dies when they come with a diagram.
The layer a search box can’t provide
The search box won the retrieval argument. The argument was never the whole job. If you are building memory for an agent whose past beliefs matter — a support agent that must not quote the plan the customer left three weeks ago, a coding agent whose confidence must survive a correction, a system whose decisions will be audited — the raw record retrieves, and the layer above it resolves. That layer is where structure finally earns its keep: supersession with a timeline, authority with provenance, retirement with history. ReFind redrew the map of what to build for retrieval; it did not redraw the map of what memory is.
OpenZync is a self-hostable starting point for that layer: a temporal knowledge graph where facts supersede rather than overwrite, exposed over REST and MCP — memory built to resolve truth, not just retrieve it. Start with the memory & context docs and the openzync-mcp repository.
This is part of a series on agent memory. Read why context windows aren’t memory, then how a temporal knowledge graph is built, then the honest map of agent memory tools, then the five patterns of graph memory, then the MCP memory-server field guide, then when agents talk to each other, who remembers, then the new attacks on AI memory, then team memory and who fixes a wrong fact, then the gap between benchmark scores and agentic memory, then what happens when agents remember things that never happened.