Getting crawled by an AI system is the easy half. The harder question is whether there is anything on your page a model can lift cleanly and attribute. We scored 243 of the world’s most-visited websites on exactly that. The median score was 45 out of 100.

Nearly two-thirds of the top of the web is structurally difficult for an AI to quote — not because the content is bad, but because of how it is built.

The distribution

Band Sites Share Meaning
Strong (80+) 8 3.3% Liftable answer, clear boundaries, attributable
Workable (55–79) 78 32.1% Extractable in parts, one or two fixes short
Weak (below 55) 157 64.6% A model would struggle to quote without distorting

Only eight sites out of 243 scored above 80. These are among the most resourced websites in existence.

Where the web fails

Pass rate by check (n=243)

Author and dates declared
5.8%
Question-form headings
16.9%
Direct answer near the top
20.6%
Cites external sources
44%
Lists or tables present
44.9%
Exactly one H1
46.1%
Structured data present
53.1%
Three or more H2 sections
58.4%
No oversized paragraphs
98.8%

Three findings worth sitting with

1. The median opening paragraph is 18 words

Against a target of 25 to 90 — long enough to contain an answer, short enough to lift whole. Sixty-nine sites had no substantial opening paragraph at all. This is the highest-weighted check in our model, and 79% of top sites fail it.

The pattern is familiar: a headline, a hero image, a short punchy line, then navigation. Beautiful for a human who chose to be there. Useless to a system scanning for a self-contained response.

2. Fewer than half have exactly one H1

H1 count Sites
Zero 105
Exactly one 112
Two 11
Three 3
Four or more 12

A twenty-year-old fundamental, failed by 54% of the biggest sites on the web. 105 have no H1 whatsoever.

3. Only 5.8% declare both an author and dates

The weakest result in the study. Retrieval systems that cite prefer content they can attribute and date. Almost none of the top of the web makes that easy, which suggests the gap is not competitive difficulty — it is that nobody has bothered.

What this means practically

If you have been told your AI visibility problem is authority, that may be true. But the data suggests a large share of pages are failing before authority is even assessed — there is simply no clean passage to extract.

The fixes are unglamorous and cheap. Open with a real answer of 25 to 90 words. Use one H1. Break content into labelled sections. Add a list or a table. Declare who wrote it and when. Link to your sources. None of that requires a budget, and 96.7% of the top of the web has not done it.

Methodology

Domains were drawn from the Tranco top-1M list (26 July 2026), with infrastructure, CDN and DNS domains excluded by pattern. We fetched each homepage with a standard browser user-agent, stripped navigation, header, footer, script and form markup, then analysed the remaining content against nine weighted checks totalling 110 raw points, normalised to 100.

The checks: direct answer near the top (20), lists or tables (15), three or more H2 sections (15), exactly one H1 (10), question-form headings (10), no paragraphs over 120 words (10), structured data present (10), author and dates declared (10), cites at least two external domains (10).

506 domains were attempted; 243 returned a fetchable homepage and were scored.

Limitations

  • These are homepages, not articles. A homepage is not primarily an answer page, so these scores understate what the same sites achieve on editorial content. The finding is best read as: if someone asks an AI about this company, can it extract a clean answer from the front door?
  • The yield rate was 48%. Homepages block automated fetches far more aggressively than robots.txt does. The scored set therefore skews toward sites with lighter bot protection.
  • The weights are ours. They reflect our judgement about what matters for extraction, not a published standard. The raw per-check data is in the dataset so you can weight it differently.
  • We caught a normalisation error mid-analysis. The checks sum to 110, not 100, and our first pass reported raw scores as percentages. Every figure here is normalised. We mention it because the value of this work depends on the numbers being right.

The data

Published openly — 506 rows with per-check results, paragraph lengths, heading counts and external link counts.

Download the dataset (CSV)

Score your own page

The same nine checks are available as a free tool. It returns the specific fix for each failure and shows your heading outline as a machine sees it.

Run the Answer Extractability Checker

This is the second study in a series. The first looked at whether AI crawlers can reach you at all. Together they cover the two halves of the problem: access, and whether there is anything worth taking once you are in.

This is part of an ongoing measurement programme. See all studies, datasets, methodology and our corrections log on the MoxSEO Research page.