Corrections
Every error we have found in our own research
We publish original data, which means we will sometimes get it wrong. When that happens the fix goes here: what broke, how we caught it, and what changed as a result.
3Logged to date
2Runs discarded
1Scoring error
0Reached a published figure
The log
What went wrong, and what changed
Two runs were discarded before release and one scoring error was caught mid-analysis. We log them anyway, because a methodology page without a corrections log is marketing — and because the failure modes below are ones anyone running similar scans will hit.
2026-07-26Discarded
High request concurrency silently corrupted a full scan
Context
The first run of the AI Crawler Blocking Study fetched robots.txt at 175 parallel requests.
What went wrong
At that concurrency google.com returned as unreachable. It fetches normally at 25 to 40. The failures were ours, not the sites’ — and they were silent. Nothing errored; rows simply came back empty.
How we caught it
A reviewer questioned why a domain as robust as google.com would be unreachable. That single implausible row exposed the whole run.
What changed
The entire first pass was discarded and rerun at 25 to 40 concurrent requests. Concurrency is now capped and stated in the methodology. No figure from the corrupted run was published.
2026-07-26Corrected
Raw scores were reported as percentages
Context
The Answer Extractability Study scores pages against nine weighted checks.
What went wrong
Those checks total 110 raw points, not 100. The first analysis pass treated the raw total as a percentage, inflating every score by roughly ten percent.
How we caught it
Caught mid-analysis when a page scored above 100.
What changed
All scores are normalised to 100 before publication, and the raw per-check values ship in the dataset so anyone can verify the normalisation. The error never reached a live page.
2026-07-26Discarded
The sample was not a list of websites
Context
Both Tranco-based studies draw from the Tranco top-1M list.
What went wrong
Tranco’s highest ranks are dominated by DNS roots, CDN endpoints and asset hosts. These serve no website and therefore no robots.txt. An unfiltered scan would have produced the headline that most top sites have no robots.txt — which is untrue.
How we caught it
Noticed when the unreachable rate looked implausibly high among the highest-ranked domains.
What changed
Infrastructure, CDN and DNS domains are now excluded by pattern before any scan begins. The exclusion is documented and reflected in every published count.
Scope
What counts as a correction
Logged here
- Anything that would have changed a published number, or did
- Discarded analysis runs
- Scoring and normalisation errors
- Sampling problems
- Figures we later found we could not evidence
Not logged here
- Typos and wording changes
- Design and layout revisions
- Adding context to an existing finding
If it does not change what a number says, it is an edit, not a correction.
Found something?
If a figure does not reconcile, tell us
Every study ships with per-row, per-check data precisely so this is possible. If a number in one of our studies does not match the dataset, we will check it and log the outcome here with the same detail as our own errors.