Skip to main content

MoxSEO

Methodology

How we run and report our studies

Every figure we publish comes from a scan we ran ourselves. This page sets out how those scans work, what the samples are, what the scoring means and where the limits sit — so you can judge the numbers rather than take them on trust.

3Published studies
1,499Domains scanned
9Weighted checks
CC BY 4.0Dataset licence
Sampling

Where the domains come from

Two of the three studies draw from the Tranco top-1M list, downloaded 26 July 2026. Tranco is a peer-reviewed ranking designed to resist the manipulation that affects commercial popularity lists, and has been used in more than 350 academic studies.

The city study samples differently. Businesses were identified per city through Google Places across agencies, development firms and local verticals, at six to nine sites per city. Domains that failed to resolve were excluded rather than substituted.

Why we filter the list first

Tranco’s highest ranks are dominated by DNS roots, CDN endpoints and asset hosts. These serve no website, and therefore no robots.txt.

Scanning the raw list would produce the headline that most top sites have no robots.txt — which is simply untrue. Infrastructure, CDN and DNS domains are excluded by pattern before any scan begins.

Scoring

The nine extractability checks

CheckWeight
Direct answer near the top20
Lists or tables present15
Three or more H2 sections15
Exactly one H110
Question-form headings10
No paragraph over 120 words10
Structured data present10
Author and dates declared10
Cites at least two external domains10
Raw total110

Two studies score pages against the same nine weighted checks, totalling 110 raw points, normalised to 100.

The weights are our judgement about what matters when a machine tries to lift a clean answer from a page. They are not a published standard, and we say so on every study that uses them.

Per-check results are published in every dataset, so you can drop a check, reweight the rest, and rerun the analysis against your own priorities.

Collection

How pages are fetched and parsed

Fetch settings

  • Standard browser user-agent. We do not disguise the scanner as Googlebot or any AI crawler. We measure what a site returns to an ordinary client.
  • Five-second timeout, three redirects. Slow or looping responses are recorded as unreachable, not retried until they succeed.
  • 25 to 40 concurrent requests. Deliberately low — higher concurrency silently corrupts results.
  • Chrome stripped before analysis. Navigation, header, footer, script and form markup removed, so scores reflect the page body.

robots.txt parsing

Files are parsed into user-agent groups, honouring consecutive User-agent lines as one group. For each of the 14 AI crawlers we track, the most specific matching group is resolved — exact name first, falling back to *.

  • BlockedDisallow: / with no Allow exceptions
  • Partial — some paths disallowed
  • Allowed — everything else

Parsing follows RFC 9309. Structured data checks follow Schema.org; heading expectations follow the HTML Living Standard.

Sample sizes

What was attempted, and what came back

StudyAttemptedUsableYield
AI Crawler Blocking
reachable robots.txt
84550659.9%
Answer Extractability
fetchable homepage
50624348.0%
City Market
20 cities, 6–9 per city
148148

Yield is a finding, not a footnote. Homepages block automated fetches far more aggressively than robots.txt does. A 48% yield means the scored set skews toward sites with lighter bot protection, and every extractability figure should be read with that in mind.

Limits

What these studies cannot tell you

Policy, not enforcement

robots.txt records what a site declares. Whether a crawler obeys it is a separate question we do not test.

Homepages, not articles

Scores understate what the same sites achieve on editorial content. Read them as: can an assistant extract a clean answer from the front door?

Small city samples

Six to nine sites shows the shape of a market, not a population figure. City percentages are indicative only.

Selection is not random

Places rankings favour businesses with strong review profiles, which likely biases the city sample upward.

Extraction is not citation

A page scoring well is easier to quote. We have not yet tested whether that correlates with actually being cited.

Our weights, our judgement

The nine checks reflect what we think matters. Per-check data is published so you can disagree using our own numbers.

Reuse

Take the data. Disagree with us using it.

All datasets are published under CC BY 4.0 — reuse, reanalyse and redistribute, including commercially, with attribution and a link. Every study ships per-row, per-check results, not just the summary figures we chose to highlight.

We have discarded two full analysis runs and corrected a scoring error mid-study. Each is documented, with what went wrong and what changed. A methodology page without a corrections log is marketing.