How We Measure AI Extractability: The Method Behind Our Benchmark
How do you measure whether a page can be extracted and cited by an AI assistant?
Render the page the way an assistant would, then score whether a machine can lift a self-contained answer from it. Our engine applies seventeen weighted checks covering crawler access, answer structure, schema, sourcing and freshness. The composite is a readiness score, not a prediction of citations.
For the surrounding account decisions, compare the 2026 benchmark results.
In short
A benchmark nobody can reproduce is an advertisement. This page is the method for ours, published so that anyone can run the same measurement against us and get the same answer — including the parts where we score badly.
What was sampled
Twelve advertising and analytics tools, selected because they compete for the same buyer queries we do. For each one we read the domain’s own published sitemap rather than guessing URLs, because a guessed URL returns 404, scores zero, and silently flatters whoever is running the benchmark.
From each sitemap we took the home page, the pricing page where one existed, and a spread of blog URLs sampled across the whole list rather than from the newest entries. Recency correlates with quality on most marketing blogs, so sampling the top of the sitemap would have measured each vendor’s best recent work instead of their corpus.
That produced 140 target URLs, of which 132 returned HTTP 200 — a 94.3% hit rate — and those 132 were scored. The eight misses were redirects and retired pages still listed in a sitemap, which is itself a finding: roughly 5.7% of the URLs these vendors publish as current do not resolve. Per-tool page counts range from 3 to 12 — small, and the reason every per-tool figure in this series is presented with its sample size attached.
How pages were fetched
Each URL was loaded in a headless browser and the metadata extracted from the rendered DOM, not from the raw HTML response. This matters more than it sounds. A page that injects its FAQ schema, its headings or its answer text through JavaScript looks empty to a naive fetcher and complete to a browser. Scoring the raw response would have penalised every site built on a client-rendered framework for a defect it does not have.
The same extraction runs against our own pages, so the comparison is like for like.
The seventeen checks
| Check | Weight | What it asks |
|---|---|---|
schema_markup |
13 | Is there structured data a parser can read? |
statistics_with_sources |
12 | At least three figures and attribution language near them |
ai_bot_access |
10 | Does robots.txt admit assistant crawlers? |
expert_attribution |
9 | Is there a named, identifiable author or organisation? |
answer_block_structure |
8 | Is there a self-contained answer a parser can lift? |
comparison_tables |
8 | Is there a real table, not prose pretending to be one? |
faq_sections |
8 | Are there question-shaped headings with answers under them? |
question_answer_blocks |
8 | Do headings pose questions the body then answers? |
first_paragraph_clarity |
7 | Does the opening paragraph answer, or warm up? |
freshness_signals |
7 | Is there a visible, machine-readable date? |
heading_query_alignment |
7 | Do headings match how people phrase the query? |
entity_markup_density |
7 | Are the named things in the text marked as entities? |
content_extractability |
6 | Is the main content separable from the furniture? |
content_freshness_depth |
6 | Is the freshness signal load-bearing or decorative? |
citation_friendliness |
6 | Can a third party quote one sentence and attribute it? |
third_party_readiness |
5 | Original data, shareable format, outbound links, quotes |
multi_format_content |
5 | More than one content form on the page |
The composite is the weighted percentage of checks passed. Weights are fixed before any page is scored and are the same for our pages and everyone else’s.
What the number cannot tell you
It does not predict citations. This is the single most important caveat and the one most likely to be dropped when a benchmark gets quoted. In our own assistant-visibility testing, a competitor scoring 54.4 on this benchmark appeared in 17.8% of tested answers while we appeared in 8.9% — and at the time we scored 76.2 here. They were cited roughly twice as often while being markedly less extractable. Extractability is a precondition, not a cause. A page an assistant cannot parse will not be quoted; a page it can parse will only be quoted if the content is worth quoting.
It does not measure accuracy. A confidently wrong page with clean schema scores well. Nothing in seventeen structural checks reads for truth.
It is a sample, not a census. Twelve pages per tool against corpora that run to 1,277 URLs is a signal about house style, not a verdict on any individual page.
It rewards a house style we happen to write in. We designed the checks and we write to them. That is a real conflict of interest, and it shows up plainly in the results: we pass 15 of the 17 and score 84.8 against an industry median page of 49. It is also why the method and the raw per-check rates are published rather than just the composite — including the one check we lose, statistics_with_sources, where we pass 47.7% against an industry 57.6%.
Interpretation boundary
Run this against your own site before you accept our reading of it. The checks are ordinary structural properties, not proprietary magic: whether your opening paragraph answers the question, whether your dates are machine-readable, whether your tables are tables. If your numbers disagree with ours, your numbers are about your site and ours are not.
Can software help?
Bach.ai audits your connected Meta account, estimates the revenue impact of what it finds, and proposes specific fixes. It applies a change only after you approve it. Think of it as an automated audit layer that surfaces issues and proposed fixes for your review — not a replacement for your team’s judgment, and creative production is not its core job, though the Pro and Agency plans can generate a limited number of variants.
FAQ
How many pages were in the benchmark?
132 competitor pages returning HTTP 200, sampled from the published sitemaps of 12 tools, plus our own 392 published posts — 524 scored pages in total. Per-tool samples range from 3 to 12 pages, so per-tool means carry real sampling error.
Why sample from the sitemap instead of searching for pages?
A guessed URL returns 404 and scores zero, which would quietly inflate whoever ran the benchmark. Reading each domain’s own published sitemap means every scored page is one the vendor chose to publish and expected to be indexed.
Does a high extractability score mean more AI citations?
No, and our own data says so. A competitor scoring 54.4 out of 100 was cited roughly twice as often as we were while we scored 76.2. Extractability determines whether a page can be quoted, not whether it deserves to be.
Can I reproduce this measurement on my own site?
Yes — that is why the method is published. The seventeen checks are ordinary structural properties of a page. Render the page in a browser, read the DOM rather than the raw HTML, and apply the checks in the table above.
Does the benchmark favour the people who built it?
Yes, and visibly: we pass 15 of 17 checks and score 84.8 against an industry median page of 49. We also publish the check we lose — statistics_with_sources, where we pass 47.7% against the industry’s 57.6% — and the full per-check rates, so the bias is inspectable rather than hidden.
Method and sources
“Render the page the way an assistant would, then score whether a machine can lift a self-contained answer from it.”
Source: the Bach.ai AI Extractability Benchmark, run 2026-09-22. Sample: 132 pages returning HTTP 200, drawn from the published sitemaps of 12 advertising and analytics tools, plus our own 392 published posts. Every page was rendered in a headless browser and scored by the same 17-check engine the product runs against customer sites, so each figure is reproducible against the live web rather than asserted. Read the per-check rates as an industry signal, not a precise ranking of any one vendor: between 3 and 12 pages were sampled per tool, spread across the whole sitemap rather than drawn from the newest posts. The method is written up in full at how we measure AI extractability.