How we run and report our studies
Every figure we publish comes from a scan we ran ourselves. This page sets out how those scans work, what the samples are, what the scoring means and where the limits sit — so you can judge the numbers rather than take them on trust.
Where the domains come from
Two of the three studies draw from the Tranco top-1M list, downloaded 26 July 2026. Tranco is a peer-reviewed ranking designed to resist the manipulation that affects commercial popularity lists, and has been used in more than 350 academic studies.
The city study samples differently. Businesses were identified per city through Google Places across agencies, development firms and local verticals, at six to nine sites per city. Domains that failed to resolve were excluded rather than substituted.
Why we filter the list first
Tranco’s highest ranks are dominated by DNS roots, CDN endpoints and asset hosts. These serve no website, and therefore no robots.txt.
Scanning the raw list would produce the headline that most top sites have no robots.txt — which is simply untrue. Infrastructure, CDN and DNS domains are excluded by pattern before any scan begins.
The nine extractability checks
| Check | Weight |
|---|---|
| Direct answer near the top | 20 |
| Lists or tables present | 15 |
| Three or more H2 sections | 15 |
| Exactly one H1 | 10 |
| Question-form headings | 10 |
| No paragraph over 120 words | 10 |
| Structured data present | 10 |
| Author and dates declared | 10 |
| Cites at least two external domains | 10 |
| Raw total | 110 |
Two studies score pages against the same nine weighted checks, totalling 110 raw points, normalised to 100.
The weights are our judgement about what matters when a machine tries to lift a clean answer from a page. They are not a published standard, and we say so on every study that uses them.
Per-check results are published in every dataset, so you can drop a check, reweight the rest, and rerun the analysis against your own priorities.
How pages are fetched and parsed
Fetch settings
- Standard browser user-agent. We do not disguise the scanner as Googlebot or any AI crawler. We measure what a site returns to an ordinary client.
- Five-second timeout, three redirects. Slow or looping responses are recorded as unreachable, not retried until they succeed.
- 25 to 40 concurrent requests. Deliberately low — higher concurrency silently corrupts results.
- Chrome stripped before analysis. Navigation, header, footer, script and form markup removed, so scores reflect the page body.
robots.txt parsing
Files are parsed into user-agent groups, honouring consecutive User-agent lines as one group. For each of the 14 AI crawlers we track, the most specific matching group is resolved — exact name first, falling back to *.
- Blocked —
Disallow: /with noAllowexceptions - Partial — some paths disallowed
- Allowed — everything else
Parsing follows RFC 9309. Structured data checks follow Schema.org; heading expectations follow the HTML Living Standard.
What was attempted, and what came back
| Study | Attempted | Usable | Yield |
|---|---|---|---|
| AI Crawler Blocking reachable robots.txt | 845 | 506 | 59.9% |
| Answer Extractability fetchable homepage | 506 | 243 | 48.0% |
| City Market 20 cities, 6–9 per city | 148 | 148 | — |
Yield is a finding, not a footnote. Homepages block automated fetches far more aggressively than robots.txt does. A 48% yield means the scored set skews toward sites with lighter bot protection, and every extractability figure should be read with that in mind.
What these studies cannot tell you
Policy, not enforcement
robots.txt records what a site declares. Whether a crawler obeys it is a separate question we do not test.
Homepages, not articles
Scores understate what the same sites achieve on editorial content. Read them as: can an assistant extract a clean answer from the front door?
Small city samples
Six to nine sites shows the shape of a market, not a population figure. City percentages are indicative only.
Selection is not random
Places rankings favour businesses with strong review profiles, which likely biases the city sample upward.
Extraction is not citation
A page scoring well is easier to quote. We have not yet tested whether that correlates with actually being cited.
Our weights, our judgement
The nine checks reflect what we think matters. Per-check data is published so you can disagree using our own numbers.
Take the data. Disagree with us using it.
All datasets are published under CC BY 4.0 — reuse, reanalyse and redistribute, including commercially, with attribution and a link. Every study ships per-row, per-check results, not just the summary figures we chose to highlight.
We have discarded two full analysis runs and corrected a scoring error mid-study. Each is documented, with what went wrong and what changed. A methodology page without a corrections log is marketing.