OpenZync
Menu

Sign In
Back to blog
engineeringAugust 18, 2026OpenZync Team22 min read

Everyone Ships an llms.txt. Almost No Agent Reads It.

Share:
On this page

The audience that never arrived

Twenty-eight percent of the web now ships an llms.txt. Ninety-seven percent of those files have never been fetched by anyone — not by a chatbot, not by a training crawler, not by an AI agent doing anything at all. The convention was proposed as robots.txt for AI, adoption swept through the industry in under two years, and the readers never came. Everyone built for an audience that does not exist.

Put a face on it, because the abstraction hides how ordinary the miss is. A team launches a marketing site in early 2026, and the launch checklist says “add llms.txt” the way an older generation of checklists said “add meta keywords.” Someone writes one — project name, a summary, a tidy list of links — and deploys it alongside the sitemap. Their site theme has started bundling the file by default, which means thousands of domains now carry one nobody wrote on purpose. The file sits at the root of the domain, well-formed and sincere, while the server logs record the truth: nothing has ever requested it. Not once.

The numbers behind that anecdote come from industry research by Ahrefs, and they are unambiguous: nearly every llms.txt file on the web has never been fetched — not once, by anything. The format did not fail loudly. It failed the quietest way a standard can fail: everyone adopted it, and nothing consumed it.

Thousands of sites ship well-formed llms.txt files at their domain roots while server logs stay silent — meanwhile an AI agent bypasses the file entirely, parsing rendered HTML, walking the accessibility tree, and calling APIs directly; everyone built the file, almost no agent came to read it

This post walks the distance between the promise and the logs: what llms.txt was supposed to be, what the largest dataset on its readership shows, why the robots.txt analogy broke, what grew in the vacuum, and what agents demonstrably read instead. The synthesis is the split this series keeps landing on — control, discoverability, and interface are three different jobs — and the interesting question was never which file to write. It was which layer to build.

What llms.txt promised

The spec is short and sincere. Jeremy Howard published it on September 3, 2024 at llmstxt.org under Answer.AI’s banner — a v2 revision followed on August 10, 2026 — and the format is exactly what a generous engineer would design for a machine reader: an H1 carrying the project name, a blockquote summarizing the project, then optional sections of markdown link lists under H2 headers. The best face of an entire site, flattened into one small file a language model could ingest without rendering anything.

The framing did the selling. “Robots.txt for AI” is a sentence every engineer already knows how to act on: one conventional filename at the domain root, one convention the whole industry has honored for decades, zero new infrastructure. Where robots.txt told crawlers where not to go, llms.txt would tell models where to look — the curated entrance instead of the crawled mess. It arrived at the exact moment site owners were panicking about AI traffic replacing search traffic, which explains why adoption outran every documentation convention before it.

Real organizations shipped it for real reasons. Stripe and Answer.AI ship one; Anthropic keeps an equivalent on its docs domain. These are not cargo-cult adopters — they are teams with the resources to maintain any convention they choose, and they chose one that costs almost nothing to maintain. Hold onto that pattern, because it is the key to reading everything that follows: the file’s production cost is genuinely near zero, which is exactly why its adoption curve says so little about its usefulness.

Worth noting too what the spec itself claimed, because the gap between claim and reception is where the confusion started. The proposal asked sites to publish a curated index and hoped model providers would encourage agents to consult it — a convention offered to the ecosystem, not a protocol demanded of it. The ecosystem took the first half and ran: agencies added the file to every deliverable, directories launched to rank the files, consultants began auditing for its presence. The second half — anyone on the reading side caring — never materialized, and no amount of publishing enthusiasm could substitute for it.

And the adoption curve looked like validation — the share of sites carrying the file climbed steadily for two years straight. What almost none of the adopters checked was the other side of the ledger: whether anyone was asking for the file. The largest dataset on that question arrived in June 2026.

The data

The Ahrefs research — Louise Linehan, with Xibeijia Guan, published June 15, 2026 — is the largest look at llms.txt readership anyone has run: 137,210 domains, the files found on them, and the server-side requests against those files. The adoption number comes with a flag Ahrefs itself attached: 28% is an upper bound, padded by parked domains and template defaults — call it roughly 38,000 domains at most. The readership number needs no flag. 97% of those files have received zero requests, ever.

The fetches that do happen are not what the format imagined. Of the requests llms.txt files receive, 96% come from bots; only 19.5% of fetchers are named AI bots; and AI retrieval bots — the systems the file was written for — account for just 1.1% of fetches. The individual names are their own commentary: GPTBot leads, Claude Code is the second-most-frequent fetcher, Slackbot out-fetched PerplexityBot, and Lighthouse audits account for roughly one fetch in a thousand. The file built for agents is mostly fetched by audit tooling and chat platforms’ link expanders.

The measurement is worth trusting precisely because of where it sits. These are server-side request logs across a very large domain sample — not surveys, not self-reporting, not extrapolation from a single site’s analytics. A file that gets fetched leaves a log line; a file that is never fetched leaves silence; and the silence here is overwhelming. The same logs that rank GPTBot first also bound how much any single fetcher could matter. Whatever noise lives in the adoption estimates, there is no version of this data in which an audience arrives late — the logs reach back far enough that a delayed readership would already be visible, and it is not.

The sharpest finding is an absence. Zero AI bots probed for the file where it was missing — no agent even looked for llms.txt on sites that don’t have one. A convention that mattered would leave traces in the negative space: agents would probe for it the way they check robots.txt before crawling. They don’t check. Google’s John Mueller put the provider-side verdict on the record, via Search Engine Land, in June 2026:

“it’s purely speculative for now (the file has existed for years, yet none of the AI systems use it)”

Trajectory points the same way. An earlier SEO-platform study from SERanking, published November 7, 2025 by Yulia Deda, measured 10.13% adoption across roughly 300,000 domains — supply rising fast, eight months before the readership data existed. Present the two findings as what they are: different companies, different dates, different methodologies, both attributed, not averaged. The shape they draw together is unmistakable — adoption climbing toward a quarter of the web while readership stayed flat at zero.

What would change the verdict is not more adoption but a reader. If any AI provider shipped an agent that requested llms.txt before rendering a page — the way every crawler requests robots.txt before crawling — the server logs would show it within weeks, and the file costs nothing to keep around in case that day comes. Until such a reader exists, the burden of proof sits with the convention rather than its skeptics, and two independent company datasets have now weighed it the same way. The claim worth evaluating is no longer whether agents might read the file someday. It is whether anything on the reading side has moved at all since the format launched.

Why it failed

Four failures compounded, and none of them is a mystery in hindsight.

No provider ever committed to reading it. robots.txt works because the companies being regulated by it are the same companies that benefit from an orderly web, and every major crawler honors it. For llms.txt, the AI labs — the only parties who could have made the file load-bearing — never shipped a parser that privileged it. Mueller’s statement is the public version of that refusal; the crawler logs are the operational one. A standard nobody is obligated to follow is a suggestion, and suggestions do not become infrastructure.

The robots.txt analogy breaks in the direction that matters. robots.txt excludes with teeth: honoring it is cheap, ignoring it invites blocks, rate-limits, and public callouts — and increasingly, regulatory attention. llms.txt suggests with none of that: no agent is required to read it, none promised to, and nothing happens when they don’t. One convention is enforced by the very parties it constrains; the other is enforced by no one. The symmetry of the filenames hid the asymmetry of the enforcement.

Agents already parse pages directly. The premise — models need a simplified digest because real HTML is too messy — aged out within a year. Frontier models read rendered pages, semantic markup, and accessibility trees natively, and the sites motivated to curate are precisely the ones whose HTML is already clean. The digest solved a problem the readers had stopped having, which is one more reason the readers never showed up.

It became a compliance ritual. Launch checklists, theme defaults, agency audits — the file spread through the same machinery that spreads meta keywords and XML sitemaps: cheap to add, easy to verify, impossible to measure. A checkbox whose completion cannot be falsified is the perfect artifact for a ritual. Nobody can prove their llms.txt failed, because nothing was ever supposed to fetch it.

There is a fifth failure underneath the four, and it is the one this series keeps meeting in different costumes: the format optimized for the producer side of the transaction because the producer side was the only side that could be organized. Site owners can be listed, audited, and sold to; agent operators cannot. Everything the ecosystem built — checklists, defaults, scorecards — addressed the people who write files, and nothing addressed the systems that were supposed to read them. A standard with a supply side and no demand side is not a standard. It is an inventory.

The GEO noise economy

Into the vacuum between “everyone ships it” and “nobody reads it,” an advice industry moved in. Ahrefs’ trends data quantifies the demand side: searches for “generative engine optimization” are up 997% over eighteen months, “geo vs seo” up 982%, and Google Cloud’s Addy Osmani redefined AEO as “Agentic Engine Optimization” — a breakout keyword at around eighty monthly searches. The fear underneath the demand is measurable too: Ahrefs estimates AI Overviews cut publisher clicks by about 58%, and queries for “zero-click search strategy” keep climbing. When clicks vanish and a new three-letter acronym trends, courses follow.

Interest versus evidence in the generative engine optimization economy: search demand for GEO tactics climbs steeply upward over eighteen months while the only controlled measurement — a single-page experiment where applying the full tactic checklist lowered citation rates — runs in the opposite direction below it

Against that demand curve, the evidence side is one controlled test, and it deserves exact scoping. Jan-Willem Bobbink ran a single-page experiment across three AI engines and posted it on LinkedIn — self-published, one page, no institution behind it. The result ran backwards: a baseline page was cited 13.34% of the time on GPT-4o-mini; after applying the full GEO checklist (“adding statistics, quotations, source citations, and an authoritative tone”), the citation rate fell to between 10.92% and 12.21%. The tactics lost to doing nothing.

Scope that honestly: one page, one model tier, self-published — it proves nothing on its own. But it is directionally consistent with what Google’s own guidance has said all along, that nothing beats being the best answer, and it sits directly against a surge in tactic demand approaching ten-fold. That gap — interest compounding while the only controlled measurement runs negative — is the noise economy: courses, audits, and “agentic seo” scorecards sold into a market with no instrument.

The dynamics around that gap matter as much as its size. A negative result on one page cannot outrun a positive curriculum at scale: courses teach the checklist because the checklist is teachable, audits sell the checklist because audits are billable, and neither has an incentive to publicize the one controlled test that ran backwards. Note also what the experiment’s design implies — the tactics were applied correctly, by an experienced practitioner, against engines actively crawling the page. This was not a strawman failure. The strongest available evidence for the tactic stack is a measurement of the tactic stack losing, and the market’s response has been to price the stack harder, not to measure it again.

Hold the honest middle, because this series does not do absolutism: some practices genuinely matter — being crawlable, being structured, being quotable — and naming them precisely is the next section’s job. The GEO industry’s problem is not that everything it sells is false. It is that the parts that work were never secret, and the parts that are secret have no evidence.

What agents actually read

The crawler logs answer “what do AI agents read” more precisely than any guide does, because logs record behavior instead of intentions. Four things show up repeatedly. None of them is named llms.txt.

Clean semantic HTML. The unglamorous baseline turns out to be the substrate: heading hierarchy that nests properly, real lists, tables that parse, pages that render without JavaScript gymnastics. The agents fetching pages are reading what users read, and the sites easiest for agents to consume are the ones already built to standards that predate the hype by a decade. Most of the “optimize for agents” advice reduces to: have good HTML.

The accessibility tree. The structured view browsers build for assistive technology — roles, names, states, relationships — doubles as a ready-made interface for agents, machine-readable by construction. Chrome’s own agentic-browsing work leans on it, and sites with sound accessibility semantics get agent-readable structure for free. The discipline your frontend team adopted for screen readers turns out to be the same discipline agents reward.

Markdown mirrors. Cloudflare — which would know, since Matthew Prince reported on June 3, 2026 that agentic traffic had passed half of all HTTP requests, a measured split of 57.5% bot to 42.5% human (reported via Tom’s Hardware the next day) — launched “Markdown for Agents” on February 12, 2026: serve a markdown rendering when the request carries Accept: text/markdown, with Cloudflare claiming an 80% reduction in token usage, in beta on Pro, Business, and Enterprise plans. The company watching the most agent traffic did not build a better llms.txt. It built content-type negotiation.

robots.txt User-Agent rules — for control, not discoverability. The one part of the old convention that carried over intact is exclusion: agents honor User-agent rules, and site owners increasingly use them to admit some crawlers and refuse others. Control works. Curation doesn’t. The file that tells agents where not to go is honored; the file that tells them where to look is ignored — and that asymmetry is the whole lesson of the format’s failure compressed into one line.

Walk one agent visit end to end and the reading list becomes concrete. A support agent is asked about refund policy, follows a link, and lands on a documentation page. It parses the rendered HTML — headings, lists, tables — because that is where the content lives. It consults the accessibility tree when the markup is ambiguous, because roles and labels disambiguate what raw tags leave unclear. If the domain serves markdown on request, it takes the cheaper rendering and keeps the saved tokens for reasoning. It checks robots.txt before crawling, because exclusion rules are enforced. At no point does anything probe for llms.txt — not because the agent objects to curation, but because nothing in its stack was ever taught to look.

The irony is now auditable. Chrome 150+ ships an experimental Agentic Browsing category in Lighthouse — documented on developer.chrome.com on May 5, 2026 — and its checks grade your llms.txt presence alongside your accessibility tree, on an informational basis for now. The audit ecosystem scores the file that agents themselves barely fetch, and grades the tree that agents demonstrably read. (PageSpeed Insights surfaces the category through third-party reports only, for now.) Build for the second signal, not the first.

One more connection closes the loop: everything above feeds the context window — what the parsers return is what the model sees — and as context engineering won’t fix your agent’s memory argued, the window is the session layer. Clean HTML and markdown mirrors make sessions better. They still do nothing about persistence, which is where the synthesis lands.

The synthesis: control, discoverability, interface

Three jobs hide inside the phrase “make your site readable by agents,” and they need three different tools. Control is robots.txt: deciding who may crawl what, with enforcement agents actually honor. Discoverability is structured content: semantic HTML, accessibility trees, markdown mirrors — pages that parse at low token cost. Interface is an API or MCP server: operations an agent can call instead of prose it can only summarize. Three jobs, three tools, three separate build decisions.

llms.txt tried to be all three and does none well. As control it has no teeth; as discoverability nobody reads it; as interface it offers links a human would click, not operations an agent can invoke. The file conflated jobs that only look similar from the outside, and the market resolved the conflation the way markets do — by using the tools that work and ignoring the one that doesn’t. This series drew the same split for memory in why context windows aren’t memory: layers that blur into one artifact fail at both jobs, and the boundary has to be built, not declared.

A worked example separates the three jobs faster than the definitions do. A publisher that refuses to have its reporting summarized needs control — robots.txt rules, honored and enforceable. A documentation site that wants its pages quoted accurately needs discoverability — semantic HTML, clean tables, maybe a markdown mirror. A software product that wants agents to create tickets, query accounts, or write memories needs an interface — authenticated endpoints, typed operations, an MCP server. One site can need all three, which is exactly why one file could never be the answer: the three builds share nothing except the word “agent.”

This series’ own project ships one anyway. OpenZync’s landing site serves an llms.txt (live at openzync.tech/llms.txt) — it costs nothing, and completeness at zero cost is still completeness. But the actual agent interface is REST and MCP: memory exposed as operations — add, search, supersede — not paragraphs describing the product. An agent that wants to use it does not read about it; it calls it. The shape of that interface category is mapped in the MCP memory-server field guide, and the multi-agent version of the same argument — interfaces, not documents, when agents talk to each other — is in when agents talk to each other, who remembers.

MCP architecture for agent memory: an AI agent connects through the Model Context Protocol to an MCP server that exposes memory operations over REST to a temporal knowledge graph backend — the agent invokes add, search, and supersede operations instead of parsing prose documents

It is a starting point, not a claim of completion. The claim is only that the layer agents can act on is built, not written — and that the honest division of labor leaves llms.txt as what it turned out to be: a free completeness gesture at the domain root, useful to humans and auditors, invisible to the agents it was named for.

The decision table

The guides keep asking whether to write an llms.txt. The honest answer is a table, because the choice is not one decision — it is five different jobs, each with its own build condition:

LayerBest atFails atBuild when
llms.txtA free pointer file for humans and auditorsNear-zero agent readership; no enforcementZero-cost completeness — ship it, expect nothing
GEO tacticsNarrative authority on paperMeasured effects run negative; unfalsifiable claimsSkip until evidence exists
Structured content + markdown mirrorsAgent parsing at low token costMaintenance overhead on every pageContent-heavy sites agents summarize from
robots.txt User-Agent rulesAccess control with teethSays nothing about discoverabilityEvery site, today
MCP / API interfaceAgent action on your productBuild cost; a contract to version and secureProducts agents should operate, not just cite
OpenZyncTemporal knowledge graph memory over REST + MCPNot an SEO play; no benchmark scores claimedProvenance and supersession; agents acting on memory

Read the rows as jobs, not as products. The first row is the Ahrefs finding restated: the file is free, so shipping it is rational, and expecting anything from it is not. The second row is Bobbink’s experiment generalized: tactics without measurement are narrative, and narrative priced as engineering is the noise economy. The third and fourth rows are what agents demonstrably use — parseable content and honored access rules — and they are boring by design, which is what working infrastructure looks like. The fifth row is the only one that changes what an agent can do rather than what it can read. The last row is this series’ own bet: memory as operations, with provenance and supersession preserved — and it claims no benchmark scores, because the measurement culture around agents is young enough that unclaimed numbers are worth more than invented ones.

If the table forces one sequencing decision, it is this: rows four and three first, row five when the product earns it, row one because it is free, row two only if evidence ever arrives. Access control costs an afternoon and is honored today. Structured content is ongoing editorial discipline with compounding returns across every reader, human and machine. An interface is a real engineering investment that should wait for a real use case — agents doing something with your product, not just describing it. And the pointer file ships in ten minutes precisely so that nobody confuses shipping it with having done the work.

What to watch

Five signals will mark where this lands. None of them are product announcements; all of them are evidence.

The CMA opt-out bites on June 17. The UK Competition and Markets Authority announced in early June 2026 that sites must be able to opt out of AI training and summarization, enforceable June 17, with the Gemini app conspicuously excluded. Google shipped a Search Console toggle to comply — and SEJ’s headline landed the catch: “Google Gives Sites AI Search Opt-Out, But Not The Data To Use It.” Watch whether opt-out reporting matures enough to act on, because an opt-out without data is a ritual of its own.

The Munich appeal resolves. A Munich court ruled in June 2026 that AI Overviews are Google’s own words and liable when wrong — covered by the-decoder.com on June 11, Reuters on June 12, WIRED on June 13 — and Google is appealing. Whichever way the appeal goes recalibrates how much responsibility AI answers carry, which is exactly the pressure that determines how hard publishers fight over every channel, this one included.

Search Console AI reporting grows past impressions. Impressions-only data cannot justify engineering work. The signal is query-level AI reporting — which prompts surfaced a site, at what position, with what click-through — turning “agentic seo” from a scorecard into a budget line. Until then, every ROI claim in the category is unfalsifiable by design.

Agentic-readiness audits move from informational to scored. Lighthouse’s Agentic Browsing category is informational today. If it ever becomes a ranking input, llms.txt gains enforcement overnight — by the back door, from the SEO side nobody planned. Watch Chrome’s documentation, not the blogs announcing it.

Citation-to-conversion measurement replaces mention-counting. The GEO economy runs on mentions because mentions are countable. The money question is whether an AI citation converts, and the vendor that instruments it honestly will deflate the noise economy faster than any skeptic. Measurement is the antidote to a market with no instrument.

The layer agents can act on

llms.txt asked agents to read. The agents that arrived read pages, walk accessibility trees, negotiate for markdown, and act through interfaces. If you are building for them, the layer worth building is the one they can operate. OpenZync is a self-hostable temporal knowledge graph where facts supersede rather than overwrite, exposed over REST and MCP — memory as operations an agent can call, with provenance and supersession preserved. Start with the memory & context docs and the openzync-mcp repository; the mechanics of the graph itself are in the five patterns of graph memory.

This is part of a series on agent memory. Read why context windows aren’t memory, then how a temporal knowledge graph is built, then the honest map of agent memory tools, then the five patterns of graph memory, then the MCP memory-server field guide, then when agents talk to each other, who remembers, then the new attacks on AI memory, then team memory and who fixes a wrong fact, then the gap between benchmark scores and agentic memory, then what happens when agents remember things that never happened, then your agent doesn’t need a knowledge graph, it needs a search box, then context engineering won’t fix your agent’s memory, then the post that started it all.