Open data

The AI Citation Index

157 readings · 2026-08-27 to 2026-09-02 · rebuilt 2 September 2026 · free to cite

Every morning we ask ChatGPT and Gemini the buying question behind one of our ranked lists, and record every website the engine returns as a source. This page is that log, published in full. The studies of AI citation we have read report their conclusions and their aggregates. This is the layer underneath one: the raw source sets, with the ask counts attached.

Two rules. Every rate on this page carries the number of asks behind it, because a page does not have a citation state — it has a rate, and a single negative reading is unreliable. And topelevens.com is excluded from both leaderboards: the questions are generated from our own list titles, so our own position in our own dataset is confounded. Our counts are stated at the bottom as a caveat, not as a finding.

How the question is phrased changes the answer more than anything on the page

We ask each category two ways. Once the way a search engine expects — best <category> <year> — and once the way a buyer who has already shortlisted actually types: <BrandA> vs <BrandB> vs <BrandC>: which is best in <year>?. Same category, same day, same engine.

The two framings return 390 and 390 distinct source domains respectively — the same number twice is a coincidence, not a typo, and the two sets are not the same sets — and they share only 49 of them — an overlap of about 13%. 341 domains appear only under the generic phrasing and 341 only under the comparison phrasing. Whatever these engines are doing, they are not consulting one ranked list of trusted publishers and reciting it.

ChatGPT · generic

10.7%

12 of 112 asks returned the page we asked about

ChatGPT · comparison

23.4%

36 of 154 asks — 2.2× the generic rate

Gemini · comparison

1.5%

1 of 66 asks, on the identical questions

These three rates measure one specific thing: how often the engine cited the Top 11 page the question was generated from. They are a property of this site as much as of the engines, and we publish them because the ask counts make them checkable, not because they generalise.

Most-cited sources — generic phrasing

Question asked: best <category> <year> · 112 asks · 718 source slots across 390 distinct domains

#DomainCitationsShare of slots
1techradar.com395.34%
2toolradar.com253.42%
3softwareadvice.com223.01%
4learn.g2.com182.47%
5fitsmallbusiness.com152.05%
6capterra.com141.92%
7dupple.com131.78%
8zapier.com101.37%
9kurums.com101.37%
10fractionaljobs.io81.1%
11g2.com70.96%
12forbes.com70.96%
13respan.ai60.82%
14stackfyi.com60.82%
15gartner.com60.82%
16reddit.com50.68%
17tomsguide.com50.68%
18thedigitalprojectmanager.com50.68%
19hostinger.com40.55%
20sureprompts.com40.55%

Most-cited sources — comparison phrasing

Question asked: <BrandA> vs <BrandB> vs <BrandC>: which is best in <year>? · 154 asks · 498 source slots across 390 distinct domains

#DomainCitationsShare of slots
1techradar.com183.49%
2reddit.com142.71%
3capterra.com71.36%
4erpresearch.com61.16%
5g2.com61.16%
6softwareadvice.com40.78%
7eightx.co40.78%
8stackfyi.com30.58%
9trustradius.com30.58%
10biztechscout.com30.58%
11saasstatshub.com30.58%
12en.wikipedia.org30.58%
13stackscored.com30.58%
14burklandassociates.com30.58%
15hackceleration.com30.58%
16propicked.com30.58%
17unvarnishedreviews.com30.58%
18langchain.com30.58%
19tomsguide.com30.58%
20techcxo.com30.58%

There is no oligopoly here

The single most-cited domain under the generic phrasing — techradar.com — accounts for 5.34% of source slots. The tail is very long: 390 domains share 718 slots, and the twentieth-placed domain is already down to 4 citations. Alongside the review aggregators you would expect — G2, Capterra, Software Advice, TechRadar — the same answers cite domains with no visible profile at all. In this dataset, being an established publisher is clearly not a precondition for being cited.

Method

  • Engines. ChatGPT (OpenAI Responses API, web search enabled); Gemini (google_search grounding). Both are asked in the same run, on the same day, with the same question.
  • Questions. Generated from a live Top 11 list page. The comparison question names the top three ranked entries on that page.
  • Repeats. The comparison question is asked more than once per page, because we measured that a single “not cited” reading is only about 75% reliable while a “cited” reading reproduces. Hits and asks are both recorded.
  • The two engines are not sampled equally. ChatGPT is asked the comparison question three times per page; Gemini once. Gemini was also asked the generic question until 1 September 2026 and it returned no citation of the asked-about page in 106 readings, so that ask was retired to stop spending on it. Read the Gemini rate as the less-sampled of the two.
  • Counting. One source slot is one domain appearing in one answer. A domain cited twice in the same answer counts once.
  • Failures. Rate-limited and errored calls are retried and are never counted as a negative reading. This has been true since 1 September 2026. Readings before that date predate retry handling, so some of them may record a call that quietly failed as a page that was not cited. Every rate here is therefore a floor.

What this dataset cannot tell you

  • It is not a sample of what buyers ask. Every question is generated from one of our own list pages, so the category mix is our category mix — business software and fractional executive services, weighted the way our site is weighted.
  • It is young. 157 readings starting 2026-08-27. The leaderboard tail is thin and individual domains near the bottom are close to noise. Treat counts under about five as indicative only.
  • It measures citation, not traffic. Being cited is not the same as being visited, and we make no claim about the second.
  • Our own numbers are confounded and we exclude them from the leaderboards. For the record: 12 generic and 18 comparison appearances for topelevens.com across the same readings. Reported separately and never ranked against the others. The questions are built from this site's own list titles, so a high count here measures the question, not the site.

Use it

The full dataset behind this page is served as JSON at /api/ai-citations, rebuilt whenever the readings are. It is free to cite, quote and republish with attribution to this page. If a figure here looks wrong, we would rather know: tell us and we will publish the correction.

Related: our ranking methodology and what we publish for AI agents.