OpenZync
Menu

Sign In
Back to blog
engineeringAugust 17, 2026OpenZync Team22 min read

Context Engineering Won't Fix Your Agent's Memory.

Share:
On this page

The skill that replaced prompt engineering

Context engineering is 2026's hottest AI skill, and it will not fix your agent's memory. Not because it is weak — because it is not that kind of tool. It is the discipline that optimizes what the model sees now. Memory is what survives across time. The two get sold as one thing, and the largest empirical study of context engineering yet — 9,649 experiments — makes the distance between them measurable.

Every guide published this year repeats the same narrative: context engineering replaced prompt engineering. The market agrees with them. Ahrefs' data from July 2026 shows searches for “ai search tracking” up 184% and “ai rank tracking” up 175% — the SEO vocabulary of an industry reorganizing itself around context. Cloudflare's CEO added a number in June 2026: agentic traffic passed 50% of internet traffic for the first time. And the model vendors keep answering the demand — Claude Opus 4.6 ships a 1M-token context window, in beta from February 2026 and generally available in March.

Put a face on it, because the abstraction hides how ordinary the gap is. A coding agent with a beautifully engineered window — the schema files laid out in the winning structure, the tool definitions pruned to the minimum, the retrieved material ranked by a context router — fixes the bug in front of it faster than last quarter's agent did. Then the ticket closes, the window closes, and the correction evaporates. Next month, the same agent hits the same failure, because nothing about the window carried the fix across time. The team that spent the year on context engineering has a faster session. It does not have an agent that learns.

The study, an arXiv preprint from February 2026, went looking for what wins inside that window — formats, file layouts, schemas. It found the format barely matters and the model is the lever. That result is a shock to the guide industry and a confirmation to anyone who has watched an agent forget: you can engineer context perfectly and the agent still forgets, still serves stale facts, still has no record of what it believed. A bigger window changes what the model can see at one moment. It changes nothing about what the agent carries across time.

Two-panel comparison of context versus memory: on the left, the context window is a bounded stack of system prompt, tool definitions, retrieved documents, and conversation history that is recomputed every call, marked with a red X badge and labeled this is your memory; on the right, agent memory persists across time as recorded past sessions, superseded facts with old truth kept and marked stale, and authority and provenance labels, with an arrow flowing from past to present; a red not-equals badge between the panels shows the context window is not memory

This post holds the honest middle, because this series does not do absolutism: context engineering is necessary, it is valuable, and it is categorically not memory. What follows is what the study found, why context and memory are different jobs, what context engineering genuinely buys, and where the two layers meet. The synthesis is the same shape this series has argued since why context windows aren't memory: engineer the context window as the session layer, buy memory as the persistence layer, and let memory feed context.

What context engineering is

The term is a Karpathy-era coinage: the argument that the prompt itself was the least interesting part of prompting — that what surrounds it, the files, the tools, the retrieved material, decides what the model can do. The discipline grew up in 2025, and its canonical statement is Anthropic's engineering guide, effective context engineering for AI agents, published September 29, 2025:

“Context engineering refers to the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference.”

The guide's definition of good work is exact: “good context engineering means finding the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome.” Note what that definition optimizes: the smallest possible set of high-signal tokens, now. There is no time axis in the definition. Nothing in it survives the end of the call.

Read the guide's own vocabulary and the discipline's boundary is visible in its failure modes. It warns about “needle-in-a-haystack style benchmarking” — the test where one relevant item sits inside a haystack of tokens and the model misses it anyway — and it names the decay that follows: “context rot,” the slow erosion of a window's usefulness as a session outgrows it. Both are within-session problems. The haystack is assembled fresh on every call, and the rot is the session's own accumulation. Neither is the forgetting that happens when the session ends, because neither is allowed to survive it.

What the discipline covers is broader than the old prompt engineering ever was. Prompt engineering optimized the words inside the prompt; context engineering optimizes everything around them — the system instructions, the tool definitions, the MCP resources, the retrieved documents, and the conversation history that get assembled into the window. That inventory is Aurimas Griciūnas' framing in his “State of Context Engineering in 2026” survey (March 22, 2026, a newsletter — a secondary source, but the most complete map of the category), and it is the honest reason the field renamed itself: the prompt stopped being the only lever, so the name had to widen.

The survey catalogs five patterns under which the 2026 practice falls: Pattern 1: Progressive Disclosure and Agent Skills; Pattern 2: Context Compression; Pattern 3: Context Routing; Pattern 4: Retrieval Evolution; and Pattern 5: Tool and Capability Management. Read the list and one thing is immediately true: every one of the five is a strategy for what enters the window at a given moment. None of them is a strategy for what the agent carries when the window is gone.

The largest study yet

The empirical test of that discipline arrived in February 2026. “Structured Context Engineering for File-Native Agentic Systems: Evaluating Schema Accuracy, Format Effectiveness, and Multi-File Navigation at Scale” is an arXiv preprint from February 2026 (arXiv:2602.05447) — a single-author preprint from Damon McMillan, submitted February 5 and revised February 12. It is 8 pages, 8 figures, 10 tables, and 26 references. No venue badge, no conference stamp: it survives or dies on its own numbers.

The scale is what makes it the largest study of its kind. The paper ran “9,649 experiments across 11 models, 4 formats (YAML, Markdown, JSON, and Token-Oriented Object Notation [TOON]), and schemas ranging from 10 to 10,000 tables,” using SQL generation as a proxy for programmatic agent operations — an agent reading schema files and writing queries against them, the exact shape of a file-native coding agent's work.

The study is careful in exactly the places a hype cycle is not. The tiers are reported separately, because the headline numbers hide opposite signs — the same file layout is a positive for frontier models and a negative for open source ones, and burying that in an average would have manufactured a finding. And the conclusion declines the move every guide makes: it does not crown a format, it tells you to tailor the architecture to the model. A preprint with that discipline does not need a venue badge to be taken seriously; it needs to be read the way it was written.

The findings, in the paper's own words:

The model is the lever. “a 21 percentage point accuracy gap between frontier and open source tiers that dwarfs any format or architecture effect.” Twenty-one points between choosing a strong model and choosing a weak one. For comparison, the largest format effect in the study is noise.

File-based context helps only the strong. “file-based context retrieval improves accuracy for frontier-tier models (Claude, GPT, Gemini; +2.7%, p=0.029) but shows mixed results for open source models (aggregate -7.7%, p < 0.001).” The same file layout that helps a frontier model measurably hurts an open source one — which is why the paper's conclusion warns against universal best practices.

Format does not matter. “format does not significantly affect aggregate accuracy” — “chi-squared=2.45, p=0.484.” YAML, Markdown, JSON, TOON: statistically indistinguishable once the model is held fixed. The format wars of the last two years of context engineering guides were fought over a null result.

There is a grep tax. The paper's abstract is careful: “compact or novel formats can incur a token overhead driven by grep output density and pattern unfamiliarity, with the magnitude depending on model capability.” A format that is dense on disk can be expensive in tokens when the agent greps it, and novel formats cost extra while the model is still learning to read them.

Ten thousand tables are navigable. “file-native agents scale to 10,000 tables through domain-partitioned schemas” — structure is what makes scale possible, which is worth holding onto for the next section.

The conclusion is the sentence to sit with: “demonstrating that architectural decisions should be tailored to model capability rather than assuming universal best practices.”

A horizontal bar chart of effect sizes from a 9,649-experiment study: model choice (frontier versus open-source) shows a large positive effect of plus 21 points, file-based context shows plus 2.7 percent on frontier models and minus 7.7 percent on open-source models, and format choice shows no significant effect with p equals 0.484 — the model is the dominant lever, not the format or the file layout

Read the chart the way the paper does: the model you choose moves accuracy by twenty-one points; the format you choose moves it by zero. The guide industry has been selling format moves for a year. The study says the only move that matters is the one the guides treat as a given.

What the study doesn't say

Read the study as narrowly as the paper does, or the second-order takes will be wrong in the same way the first-order ones were. The task is SQL generation against file-native schemas — structured data, well-defined schemas, a proxy for programmatic operations. It is not conversational context, not retrieval strategy, and not memory. The paper says plainly what it measured; the discipline should say equally plainly what it did not.

The first over-read is “format is noise, so context engineering is dead.” Wrong twice. The study held the contents of the context fixed and varied the container — YAML, Markdown, JSON, TOON. What you put in the window was not the variable. Selection quality — which documents, which schema slices, which conversation turns — still dominates everything the format comparison could not see, because the agent can only answer from what it was given. A study that finds the bottle is irrelevant says nothing about the wine.

The second over-read is “structure is pointless.” Also wrong, and the paper's own fifth finding is the counter: file-native agents scale to 10,000 tables through domain-partitioned schemas. Structure is what makes the file layout navigable at all; the domain-partitioned schema is the reason a model can find the right table among ten thousand. What the study killed is not structure — it is the claim that one structure beats another on accuracy. The schema exists to be navigated, not to win a format contest.

The honest middle, in one sentence: this study is about which structure, not whether context matters, and it is about file layout, not retrieval strategy and not memory. The lesson it does support is the one in its conclusion — tailor the architecture to the model — which is a context engineering lesson, not a memory lesson. Teams that need memory are still waiting for a study that touches the time axis, and the study that comes closest exists in this series' other file: ReFind, from the raw-logs post, measured retrieval over the record and found it needs no structure at all — while the questions that needed a timeline went unanswered by every retrieval system in the comparison.

Put the boundary to work. “Which table holds the customer's plan?” is a file-native question: the right schema slice, correctly formatted, and the model answers from the window. “Are they still on the plan we moved them to in March?” is the same words and a different job — the answer depends on what has happened since, on supersession, on a timeline no window holds. The first question is the study's domain. The second is the reason the persistence layer exists. Teams that blur the two buy a year of context engineering and then wonder why the agent still quotes the plan the customer left.

Why context can't be memory

The categorical split is the whole argument, so state it exactly. Context is what the model sees now — bounded, ephemeral, recomputed from scratch on every call. Memory is what survives across time — the sessions, the state, the truth-changes that persist when the window closes. The two properties are orthogonal: a context window of a million tokens is still a context window. The 1M-token milestone changed the what — how much the model can see at once — and changed nothing about the when — how long anything it sees survives. Post one of this series made the point when the largest windows were a rounding error compared to today: context windows aren't memory, and the capacity arms race did not repeal that.

The evidence this series has gathered since makes the split empirical, not just definitional. ReFind, covered in the raw-logs post, is the cleanest demonstration: an agent with perfect context selection — raw logs, lexical index, an iterative search loop that decides exactly what enters its window, retrieval at 58.2 against HippoRAG 2's 53.2 — still cannot answer “is this fact true now?” The selection was flawless and the current truth stayed out of reach, because current truth is not a property of what you retrieve, it is a property of what has happened since. Context selection is not truth resolution. A bigger, better-engineered window is the same shape of answer, at higher resolution.

The authority-collapse preprint, covered in the confabulation post, adds the second gap. Context can carry a claim; it cannot carry the claim's standing — who said it, under what conditions, with what permission. Consolidation “preserves a claim while erasing the source constraints governing its authorized use,” and the benchmark showed that persisting authority labels cuts the unauthorized-action rate from 16.9% to 0.0%. The label is memory work: it has to survive the session, attach to the fact, and travel with it. A context window can carry the claim now. It cannot carry the permission across time.

And then there is the question this series keeps returning to: what did the agent believe when it acted? A wrong belief that steers a thousand later actions — the poisoning post documented the mechanism, the confabulation post the self-written variant — is not discoverable in any window, because the window shows what the model saw, not what the agent held to be true. Belief is a property of a timeline: what was stored, when, and what superseded it. Context cannot reconstruct that. Memory exists because of it.

What context engineering genuinely buys

None of that makes context engineering worthless, and the honest accounting is the point of this section. Within the session, the discipline buys four real things, each one evidenced.

Attention economics. The lost-in-the-middle result — Liu et al., “Lost in the Middle: How Language Models Use Long Contexts” (arXiv:2307.03172, TACL 2024) — showed models use the ends of long contexts and lose the middle; the needle-in-a-haystack test, Greg Kamradt's 2023 probe, showed a single relevant item buried in a haystack gets missed regardless. The mechanism is real and it is the reason the guide's definition works: every token you leave out is a token that cannot dilute attention. The “smallest possible set of high-signal tokens” is attention economics stated as engineering.

Progressive disclosure at scale. The first of the five patterns is the one with a standard now: Agent Skills, which the survey is explicit about — Agent Skills are the standard implementation of progressive disclosure. Anthropic launched Agent Skills on October 16, 2025, and published the format as an open standard on December 18, 2025, with adoption spreading to OpenAI's Codex CLI, Google's Gemini CLI, and a growing list of others. The mechanism is exactly the smallest-token-set philosophy made operational: skills are markdown files with YAML frontmatter, roughly 80 tokens at discovery, expanding to their full 275–8,000-token instruction only when activated. The tradeoffs are documented too — accuracy degrades beyond 100 skills, as overlapping descriptions cause misactivation, and the open question remains when an activated skill gets deactivated. Progressive disclosure is a context engineering win; the 80-token discovery cost versus the 8,000-token activation cost is the whole case for it.

Cost. What enters the window sets the serving bill, and the Total-Recall preprint, covered in the raw-logs post, priced it: memory serving cost versus accuracy, with break-even ranging from “the first tens of turns” for the cheapest systems to “never within 400 turns” for the most expensive. Context choice drives the cost side of that ledger harder than the memory system does, because tokens are the currency of both.

Fewer failed retrievals. ReFind's search loop is context selection in its purest form — the agent deciding, iteratively, what enters its window. The 58.2 result is a context engineering result: give the agent control over its own context assembly and retrieval precision follows. The discipline that decides what goes in is the same discipline that decides what the model can use.

Everything on that list happens inside a single session. That is the point, not a limitation smuggled in: context engineering is the session layer, and its wins are session wins. The discipline that maximizes the likelihood of an outcome now is doing exactly what it was built for — and exactly nothing about the next call.

The synthesis: engineer context, buy memory

The architecture that survives both the preprint and the series' evidence is two layers, not one. The session layer is context engineering: the smallest possible set of high-signal tokens, per Anthropic's definition, assembled per call — the tool definitions, the retrieved documents, the progressive disclosure, the compression. The persistence layer is memory: the record, the authority labels, the timeline that resolves what is true now and what the agent believed then. One optimizes the window. The other survives it.

The layers feed each other in one direction. Memory feeds context: retrieval decides which records, which facts, which past beliefs enter the window at all. Context engineering decides how they enter — formatted, compressed, disclosed progressively. The boundary this series drew in the raw-logs post stands: search and context selection retrieve; the temporal layer resolves. The context window is where memory shows up to work. It is not where memory lives.

Walk one session against that architecture. The agent opens a ticket from a customer who upgraded in March. Context engineering assembles the window: the current plan, the migration status, the billing rules, formatted and pruned to the smallest high-signal set. Memory is what put the March upgrade in the window at all — the retrieval layer found the plan record, and the temporal layer resolved that the March upgrade is the current truth, not the February plan that search also resurfaced. The window is where the work happens; the persistence layer is why the window knew what to ask for. Engineer the window well and the session is fast. Buy memory well and the session is the second visit, not the first.

A layered architecture: context engineering assembles the session window at the top from a retrieval layer below it, and a temporal memory layer sits under both, persisting what the agent believed across time — the temporal layer feeds the window while keeping the record, superseding rather than overwriting facts, and preserving authority and provenance labels across sessions

That is the direction OpenZync works in: a self-hostable temporal knowledge graph where facts supersede rather than overwrite, feeding the context window without claiming to replace it — the persistence layer, made operational. It is a starting point, not a claim of completion; where it sits among the other tools in this category is mapped honestly in the openzync-vs-zep-vs-mem0 post, and the mechanics are in how a temporal knowledge graph is built.

The decision table

The comparison guides keep asking which skill to learn first. The honest answer is a table, because the choice is not “context engineering vs memory” — it is which job each layer is being asked to do:

LayerBest atFails atBuild when
Context engineeringSmallest high-signal window for a sessionSession-scoped only; nothing survives the callIn-session precision work at the token frontier
RetrievalFinding relevant records to feed the windowStale-fact resurfacing; no timelineSearch over large records before reasoning
MemoryPersistence and truth-resolution across timeConstruction and serving costPast beliefs and audit matter
OpenZyncTemporal knowledge graph with supersede-not-overwriteRaw retrieval when search sufficesProvenance and audit; no benchmark scores claimed

Read the rows as jobs, not as products. The first row is the study's entire case: within a session, context engineering is the only lever you control — the format wars are over and the model is the floor — and the row concedes nothing across time. The second row is ReFind's case: retrieval decides what enters the window, and it needs no structure to do it. The third row is the categorical claim of this post: memory is the only row that survives the window, and it costs construction to build and tokens to serve. The last row is this series' own bet: the temporal layer is not a context upgrade; it is the layer that keeps what was true then, preserves authority, and makes belief auditable — and it claims no benchmark scores for the retrieval job, because the preprints above just redrew that map. Choose the row by the question the agent actually has to answer, and the trend lists start answering themselves.

What to watch

Five signals will mark where the category lands. None of them are product announcements; all of them are evidence.

Context protocols standardize. MCP is on its way to becoming the context protocol — the way tools, resources, and memory get declared to the window — and the memory-server field guide mapped the shape of that category last month. The maturity signal is when context assembly stops being bespoke prompt furniture and becomes a protocol surface: agents declaring what they need, servers declaring what they offer, the window assembled by contract.

Dynamic context assembly becomes the default. McMillan's stated future direction is agents that determine their own needed context instead of receiving a fixed assembly. The study's finding — that the same file layout helps frontier and hurts open source — is an argument for exactly that: if the optimal context depends on the model, the model's agent should be the one assembling it. The signal is when context engineering shifts from authoring to supervising: the agent builds the window, the engineer audits it.

Capability becomes the design input. The 21-point gap says the model is the lever, and the honest implication is architectural: design the context for the model tier you actually run, not the tier the guides assume. The signal is context-aware fine-tuning — models trained to work with smaller, high-signal windows instead of assuming the full haystack — and the study's “tailored to model capability” conclusion becoming a design requirement rather than a closing caveat.

The memory-context boundary becomes a named architecture split. This post is one data point; the split is real when it stops needing a post. The signal is standard architecture vocabulary: session layer and persistence layer named as distinct concerns, with context engineering hired for the first and memory for the second — the same way compute and storage stopped being one thing in database design.

Cost-aware context evaluation becomes the standard. Total-Recall priced memory at tokens per correct answer; the same discipline applies to context. The signal is context eval reporting the price of its own windows — how many tokens the winning assembly costs, not just what it scores — because the 400-turn break-even curve made the cost of a bloated window measurable, and a score without a serving cost is a score nobody can budget against.

The persistence layer a context window can't provide

Context engineering won the argument about the session. The argument was never the whole job. If you are building an agent whose behavior has to survive the window — a support agent that must not quote the plan the customer left three weeks ago, a coding agent whose corrections must persist across a thousand files, a system whose past beliefs will be audited — the window assembles, and the layer beneath it persists. That layer is where memory earns its keep: supersession with a timeline, authority with provenance, belief with history. The preprint showed what context engineering can do inside a session. It did not touch the time axis, and the time axis is the one your agent actually fails on.

OpenZync is a self-hostable starting point for that layer: a temporal knowledge graph where facts supersede rather than overwrite, exposed over REST and MCP — the persistence layer a context window can't provide, and a starting point for teams building memory, not just context. Start with the memory & context docs and the openzync-mcp repository.

This is part of a series on agent memory. Read why context windows aren't memory, then how a temporal knowledge graph is built, then the honest map of agent memory tools, then the five patterns of graph memory, then the MCP memory-server field guide, then when agents talk to each other, who remembers, then the new attacks on AI memory, then team memory and who fixes a wrong fact, then the gap between benchmark scores and agentic memory, then what happens when agents remember things that never happened, then your agent doesn't need a knowledge graph, it needs a search box.