Half of Advertising Tools Publish an llms.txt. Does It Do Anything?
Should my site publish an llms.txt file?
It is cheap and harmless, so probably yes — but do not expect it to move anything on its own. Six of twelve advertising tools we measured publish one, and their mean extractability score was within two points of the six that do not.
For the surrounding account decisions, compare the full benchmark and the method.
In short
Six of twelve tools (50%) publish an llms.txt. Madgicx, Triple Whale, Motion, Hyros, Supermetrics and Whatagraph have one; Revealbot, AdEspresso, Northbeam, Smartly.io, Adriel and Foreplay do not.
Adoption has crossed half the category. Whether it is doing anything is a separate question, and the honest answer from our data is: not much on its own.
What the file is
llms.txt is a proposed convention: a markdown file at the root of a domain that lists the pages a site considers most useful to a language model, with short descriptions. It is a curation hint, not an access control — that job belongs to robots.txt.
The distinction matters because the two get conflated. robots.txt decides whether a crawler may fetch you. llms.txt suggests what is worth reading once it has. In our sample all 132 pages passed the crawler-access check and none were bot-blocked, so access is not the constraint in this category. Curation might be.
What the data shows
The six tools that publish one score a mean of 49.7; the six that do not score 48.2. A 1.5-point gap on samples of six domains is indistinguishable from noise, and the confound is obvious: publishing llms.txt is a marker of a team that pays attention to this area, which is also correlated with everything else they do.
The file is not doing the work. It is a symptom of teams that are doing the work.
The case for publishing one anyway
It costs roughly an hour and has no downside we can find. More usefully, writing it forces a decision most sites have never made: which twenty pages actually represent us?
Most marketing sites cannot answer that. They have a corpus that grew by accretion, and the honest answer is “we do not know which pages are load-bearing”. Producing a curated list is a genuinely clarifying exercise independent of whether any consumer reads the file — it usually surfaces that the pages the team is proudest of are not the pages answering the queries people ask.
How to write one that is not useless
The failure mode is a dump of every URL, which curates nothing and is strictly worse than the sitemap that already exists.
- Keep it short. Twenty to fifty entries. If everything is important, nothing is.
- Write real descriptions. One line saying what question the page answers, not the meta description.
- Group by intent, not by category. “Choosing between tools”, “diagnosing a delivery problem”, “setting up measurement” beats “Blog”, “Guides”, “Resources”.
- Include the pages that say what you do not do. Limits and non-features are disproportionately useful to a model trying to answer whether you fit a described situation.
- Keep it current. A curation file listing retired pages is a liability. In our sample 5.7% of sitemap URLs across the category did not resolve, so this failure mode is already common in files that are generated rather than maintained.
Interpretation boundary
Six domains per group is far too small to measure an effect, and we did not test whether any assistant actually consumes the file — that would need a controlled experiment we have not run. Our finding is narrow: adoption is at 50% in this category, and the presence of the file does not predict extractability. Neither of those is evidence that it does nothing.
Can software help?
Bach.ai audits your connected Meta account, estimates the revenue impact of what it finds, and proposes specific fixes. It applies a change only after you approve it. Think of it as an automated audit layer that surfaces issues and proposed fixes for your review — not a replacement for your team’s judgment, and creative production is not its core job, though the Pro and Agency plans can generate a limited number of variants.
FAQ
How many companies publish an llms.txt file?
Six of the twelve advertising tools we measured in September 2026 — 50%. Adoption has crossed half the category, though its presence did not predict a higher extractability score.
Does llms.txt improve AI visibility?
Our data cannot show that it does. Publishers scored a mean 49.7 against 48.2 for non-publishers, a gap indistinguishable from noise on six domains per group. The file appears to be a marker of attentive teams rather than a cause.
What is the difference between llms.txt and robots.txt?
robots.txt decides whether a crawler may fetch your pages; llms.txt suggests which pages are worth reading. They are not substitutes. In our sample all 132 pages already allowed assistant crawlers, so access was never the constraint.
What makes an llms.txt file useful?
Curation. Twenty to fifty entries with real one-line descriptions of what each page answers, grouped by reader intent rather than site category, and including the pages that describe your limits. A dump of every URL is worse than the sitemap you already have.
Method and sources
“It is cheap and harmless, so probably yes — but do not expect it to move anything on its own.”
Source: the Bach.ai AI Extractability Benchmark, run 2026-09-22. Sample: 132 pages returning HTTP 200, drawn from the published sitemaps of 12 advertising and analytics tools, plus our own 392 published posts. Every page was rendered in a headless browser and scored by the same 17-check engine the product runs against customer sites, so each figure is reproducible against the live web rather than asserted. Read the per-check rates as an industry signal, not a precise ranking of any one vendor: between 3 and 12 pages were sampled per tool, spread across the whole sitemap rather than drawn from the newest posts. The method is written up in full at how we measure AI extractability.