Skip to main content

MoxSEO

Research

We measure things about the web that nobody else publishes

Most agency research is a survey of 200 marketers. Ours is instrumentation. We build the measuring tool, run it across thousands of real websites, publish the finding, and release the raw data under an open licence.

2 published studies1,351 data rows releasedCC BY 4.0 open licence3 corrections logged
The programme

AI search asks three questions of every website

Almost all published advice addresses the first and ignores the other two. We are working through all three.

01 · ACCESS

Can AI reach you?

Whether robots.txt permits the crawlers behind ChatGPT, Claude, Perplexity, Gemini and Apple Intelligence.

Published · 506 domains
02 · EXTRACTION

Can AI quote you?

Whether a page is structurally capable of yielding a clean, attributable answer a model can lift.

Published · 243 sites
03 · COMPREHENSION

Can AI understand you?

Whether structured data and entity signals let a machine classify you correctly and confidently.

In progress
Published

Studies and open datasets

1.9×
Training crawlers blocked more often than retrieval crawlers

The Web Is Blocking AI Training, Not AI Search

26 JUL 2026  ·  506 DOMAINS  ·  845 ROWS  ·  ASHISH KHAN

Two independent operators, near-identical ratios. The web has drawn a deliberate line: do not train on my content, but do cite me in answers. Blocking is overwhelmingly intentional — 95 of 104 blockers use named rules rather than blanket wildcards.

3.3%
Share of top websites structurally quotable by AI

Only 3% of Top Websites Are Structurally Quotable by AI

26 JUL 2026  ·  243 SITES  ·  506 ROWS  ·  SAKSHI KUMARI

Median extractability score of 45 out of 100. The median opening paragraph runs 18 words against a 25 to 90 word target, 54% of sites lack exactly one H1, and only 5.8% declare both an author and dates.

Method

How we work

  • Instrument first. Every study is powered by a tool we built and published, so you can check a single case against our method before trusting the aggregate.
  • Academic sampling. Domains come from Tranco, a peer-reviewed ranking designed to resist manipulation and used in over 350 studies.
  • Yield rates stated. We publish how many domains we attempted, not only how many we scored. A study that hides its denominator is not a study.
  • Open data. Raw per-domain results under CC BY 4.0, so you can weight them differently or disagree with us using our own numbers.
  • No surveys. We measure what sites actually do, not what marketers say they do.
Corrections log

Errors we caught and fixed

Research you cannot audit is marketing. Mistakes found mid-study are the strongest evidence anyone was checking.

Tranco’s top list is not a list of websites

Our first sample was dominated by DNS roots and CDN endpoints, which serve no robots.txt because they serve no site. Uncorrected, we would have reported that most top sites have no robots.txt. Caught before publication.

High concurrency silently corrupted results

At 175 parallel requests, google.com appeared unreachable. At 25 to 40 it fetches normally. An entire first pass was discarded. Caught by testing a known-good domain alone.

Extractability checks sum to 110, not 100

Our first analysis reported raw points as percentages, producing a site scoring 110 out of 100. All published figures are normalised; the tool was never affected. Caught during analysis.

Reuse

Citing this research

Both datasets are released under CC BY 4.0. Reproduce, adapt and build on them commercially, with attribution.

MoxSEO (2026). AI Crawler Access Study.
https://moxseo.com/ai-crawler-blocking-study/

MoxSEO (2026). Answer Extractability Study.
https://moxseo.com/ai-answer-extractability-study/

If you use the data somewhere published, we would like to see it — tell the team.

Roadmap

What is next

Both studies re-run quarterly, so the trend becomes visible rather than the snapshot. The third pillar — structured data comprehension across the same sample — is in progress and uses our schema validator.

Beyond that: hreflang correctness at scale, llms.txt adoption over time, and whether extractability scores correlate with actual AI citation frequency.

If you have a measurement question you would fund or collaborate on, we are open to it.

Why an agency publishes this

Anyone can claim expertise. Publishing your data and your mistakes is a different proposition.

Partly because it is genuinely useful. Mostly because an SEO agency should be judged on evidence, and this is the hardest form of evidence to fake.

AI SEO servicesFree toolsHow we work