---
title: Only 3% of Top Websites Are Structurally Quotable by AI: 243 Sites Scored
description: We scored 243 of the world's most-visited websites on whether an AI can extract a clean answer from them. The median score was 45 out of 100.
url: https://moxseo.com/ai-answer-extractability-study
date_modified: 2026-07-26
author: Sakshi Kumari
language: en_US
---

Getting crawled by an AI system is the easy half. The harder question is whether there is anything on your page a model can lift cleanly and attribute. We scored 243 of the world’s most-visited websites on exactly that. **The median score was 45 out of 100.**

Nearly two-thirds of the top of the web is structurally difficult for an AI to quote — not because the content is bad, but because of how it is built.

## The distribution

| Band | Sites | Share | Meaning |
| --- | --- | --- | --- |
| Strong (80+) | 8 | 3.3% | Liftable answer, clear boundaries, attributable |
| Workable (55–79) | 78 | 32.1% | Extractable in parts, one or two fixes short |
| Weak (below 55) | 157 | 64.6% | A model would struggle to quote without distorting |

**Only eight sites out of 243 scored above 80.** These are among the most resourced websites in existence.

## Where the web fails

Pass rate by check (n=243)

Author and dates declared

5.8%

Question-form headings

16.9%

Direct answer near the top

20.6%

Cites external sources

44%

Lists or tables present

44.9%

Exactly one H1

46.1%

Structured data present

53.1%

Three or more H2 sections

58.4%

No oversized paragraphs

98.8%

## Three findings worth sitting with

### 1. The median opening paragraph is 18 words

Against a target of 25 to 90 — long enough to contain an answer, short enough to lift whole. Sixty-nine sites had no substantial opening paragraph at all. This is the highest-weighted check in our model, and **79% of top sites fail it**.

The pattern is familiar: a headline, a hero image, a short punchy line, then navigation. Beautiful for a human who chose to be there. Useless to a system scanning for a self-contained response.

### 2. Fewer than half have exactly one H1

| H1 count | Sites |
| --- | --- |
| Zero | 105 |
| Exactly one | 112 |
| Two | 11 |
| Three | 3 |
| Four or more | 12 |

A twenty-year-old fundamental, failed by 54% of the biggest sites on the web. 105 have no H1 whatsoever.

### 3. Only 5.8% declare both an author and dates

The weakest result in the study. Retrieval systems that cite prefer content they can attribute and date. Almost none of the top of the web makes that easy, which suggests the gap is not competitive difficulty — it is that nobody has bothered.

## What this means practically

If you have been told your AI visibility problem is authority, that may be true. But the data suggests a large share of pages are failing before authority is even assessed — there is simply no clean passage to extract.

The fixes are unglamorous and cheap. Open with a real answer of 25 to 90 words. Use one H1. Break content into labelled sections. Add a list or a table. Declare who wrote it and when. Link to your sources. None of that requires a budget, and **96.7% of the top of the web has not done it**.

## Methodology

Domains were drawn from the [Tranco](https://tranco-list.eu/) top-1M list (26 July 2026), with infrastructure, CDN and DNS domains excluded by pattern. We fetched each homepage with a standard browser user-agent, stripped navigation, header, footer, script and form markup, then analysed the remaining content against nine weighted checks totalling 110 raw points, normalised to 100.

The checks: direct answer near the top (20), lists or tables (15), three or more H2 sections (15), exactly one H1 (10), question-form headings (10), no paragraphs over 120 words (10), structured data present (10), author and dates declared (10), cites at least two external domains (10).

506 domains were attempted; 243 returned a fetchable homepage and were scored.

## Limitations

- **These are homepages, not articles.** A homepage is not primarily an answer page, so these scores understate what the same sites achieve on editorial content. The finding is best read as: if someone asks an AI about this company, can it extract a clean answer from the front door?
- **The yield rate was 48%.** Homepages block automated fetches far more aggressively than robots.txt does. The scored set therefore skews toward sites with lighter bot protection.
- **The weights are ours.** They reflect our judgement about what matters for extraction, not a published standard. The raw per-check data is in the dataset so you can weight it differently.
- **We caught a normalisation error mid-analysis.** The checks sum to 110, not 100, and our first pass reported raw scores as percentages. Every figure here is normalised. We mention it because the value of this work depends on the numbers being right.

## The data

Published openly — 506 rows with per-check results, paragraph lengths, heading counts and external link counts.

[Download the dataset (CSV)](https://moxseo.com/wp-content/uploads/moxseo-research/answer-extractability-study-2026-07-26.csv)

## Score your own page

The same nine checks are available as a free tool. It returns the specific fix for each failure and shows your heading outline as a machine sees it.

[Run the Answer Extractability Checker](https://moxseo.com/tools/ai-answer-extractability-checker/)

This is the second study in a series. The first looked at [whether AI crawlers can reach you at all](https://moxseo.com/ai-crawler-blocking-study/). Together they cover the two halves of the problem: access, and whether there is anything worth taking once you are in.

This is part of an ongoing measurement programme. See all studies, datasets, methodology and our corrections log on the [MoxSEO Research](https://moxseo.com/research/) page.
