We scanned 148 businesses across 20 cities in five countries with our own extractability and crawler-access tools. Structured data is nearly universal. Heading architecture is close behind. And fewer than one site in six tells you who wrote the page or when.

That gap holds everywhere. It is not a developing-market problem or a small-business problem. New York, London and Sydney — three of the most expensive agency markets in the world — all sit at or below 17%.

The headline numbers

Pass rate by check · 148 sites, 20 cities

Three or more H2 sections
90.5%
Structured data present
85.1%
Lists or tables
75%
Exactly one H1
70.3%
Question-form headings
66.2%
Cites external sources
62.8%
Direct answer near top
41.2%
Author and dates declared
16.2%

Median extractability across the whole sample was 68 out of 100, well above the 45 median we recorded across 243 top global websites in our earlier study. These are professional service businesses — agencies, developers, clinics, hotels, jewellers — and they build better than the web average.

Attribution is the universal failure

Eight checks. Seven of them pass on most sites. One does not come close.

# City Country Author + dates
1 Chennai India 43%
2 Bengaluru India 38%
3 Kolkata India 33%
4 Ahmedabad India 33%
5 Toronto Canada 29%
6 Singapore Singapore 29%
7 Gurgaon India 22%
8 Mumbai India 17%
9 London UK 17%
10 Delhi India 13%
11 Jaipur India 13%
12 New York USA 13%
13 Pune India 11%
14 Sydney Australia 11%
15 Lucknow India 0%
16 Indore India 0%
17 Noida India 0%
18 Bhopal India 0%
19 Chandigarh India 0%
20 Dubai UAE 0%

Six cities recorded zero. Not one sampled site in Lucknow, Indore, Noida, Bhopal, Chandigarh or Dubai declares both an author and publish dates.

And the ranking runs against expectation. The top four markets on attribution are all Indian. Chennai at 43% and Bengaluru at 38% beat New York, London, Sydney and Dubai — markets with far higher median scores and vastly larger budgets.

India versus international

Region Sites Median score Author + dates
India (14 cities) 103 64 16.5%
International (6 cities) 45 73 15.6%

International markets lead on median score by nine points. On attribution they are statistically identical — 15.6% against 16.5%. Whatever drives the score gap, it is not the willingness to sign your work. Nobody does that anywhere.

The second failure: opening paragraphs

Only 41.2% of sites open with a paragraph between 25 and 90 words — long enough to contain an answer, short enough for a model to lift whole. This is the highest-weighted check in our model, because it determines whether extraction is possible before anything else is assessed.

# City Direct answer near top
1 Kolkata 67%
2 New York 63%
3 Chennai 57%
4 Ahmedabad 56%
5 Delhi 50%
6 London 50%
7 Pune 44%
8 Sydney 44%
9 Noida 43%
10 Bengaluru 38%
11 Jaipur 38%
12 Dubai 38%
13 Mumbai 33%
14 Indore 33%
15 Bhopal 33%
16 Chandigarh 33%
17 Toronto 29%
18 Singapore 29%
19 Lucknow 25%
20 Gurgaon 22%

Only two markets clear 60%: Kolkata at 67% and New York at 63%. Fourteen of twenty sit below half.

What is already solved

It is worth being clear about what these businesses do well, because the picture is not one of general neglect.

90.5% use three or more H2 sections. 85.1% carry structured data — and in New York, London, Sydney and Singapore that figure is 100%. 75% include lists or tables. These are competent, well-built websites.

The technical layer is done. The editorial layer is not.

AI crawler access

Nine of 148 sites block at least one AI crawler. Twenty-two name AI user-agents explicitly in robots.txt without blocking them. Pune had the highest deliberate engagement, with three of nine sites naming crawlers and two blocking.

That is a far lower blocking rate than the 20.6% we found among the world’s top 506 domains. Publishers block. Professional service businesses mostly do not — they want the visibility.

Why this matters

Retrieval systems that cite prefer content they can attribute. A named author and a publish date are the cheapest signals available — a template change, not a strategy. They cost nothing and almost nobody has claimed them.

If you run a professional services business, this is the rare case where the competitive gap is genuinely trivial to close. Add a byline. Add a date. Open with the answer.

Methodology

Businesses were identified per city via Google Places across agencies, development firms and relevant local verticals. Each live homepage was fetched with a standard browser user-agent and analysed against nine weighted structural checks totalling 110 raw points, normalised to 100. robots.txt was parsed separately for 14 AI crawlers.

Sample sizes ranged from 6 to 9 sites per city. Domains that did not resolve were excluded rather than substituted. Hyderabad was scanned separately for its own city page and its raw rows are not included in this pooled dataset.

Limitations

  • Small per-city samples. Six to nine sites shows the shape of a market, not a population figure. City-level percentages should be read as indicative.
  • Homepages, not articles. A homepage is not primarily an answer page, so these scores understate what the same businesses achieve on editorial content.
  • The weights are ours. They reflect our judgement about what matters for extraction, not a published standard. Per-check results are in the dataset so you can weight them differently.
  • Business selection is not random. We used Places rankings, which favour businesses with strong review profiles. This likely biases the sample upward.

The data

All 148 rows, with per-check results, opening paragraph length, heading counts and crawler verdicts.

Download the dataset (CSV)

Released under CC BY 4.0. Per-city breakdowns are published on each city market study.

Check your own site

Run the Answer Extractability Checker

If you pass the author-and-dates check, you are already ahead of 84% of the professional service businesses we measured across five countries.

Sources and standards

The checks in this study are grounded in published documentation rather than our own opinion. Structured data definitions follow Schema.org. Crawler directive parsing follows the Robots Exclusion Protocol as specified in RFC 9309. Heading and document-outline expectations follow the HTML Living Standard, and general crawling and indexing guidance follows Google Search Central.