AI knowledge base

Research · published 2026-08-19

Which sources AI answer engines cite when buyers ask about business software

When somebody asks an AI assistant which software to buy, which websites actually decide the answer?

Short answer

For commercial software buying questions, AI citations are extraordinarily fragmented, not concentrated. Across 3,952 citations we recorded 2,031 distinct domains, 71% of which were cited exactly once, and the top 25 domains together accounted for 13.7% of all citations. Reddit, the domain most often described as the single biggest source in AI answers, was 0.71%. Wikipedia was 0.18%. Capterra was zero.

The measurement

3,952 citations were recorded across 608 answers to 294 questions. They came from 2,031 distinct domains, and the top 25 of those account for 13.7% of all citations.

#DomainCitationsShareAnswers it appeared in
1salesforce.com360.91%29
2lindy.ai310.78%22
3pipeline.zoominfo.com300.76%23
4reddit.com280.71%24
5g2.com260.66%20
6monday.com260.66%21
7zapier.com260.66%25
8arxiv.org250.63%15
9hubspot.com240.61%23
10knowledge.apollo.io230.58%18
11apollo.io220.56%19
12salesforge.ai220.56%18
13agorareal.com200.51%10
14linkedin.com200.51%14
15medium.com200.51%14
16clay.com190.48%15
17aiforcrecollective.com180.46%13
18dev.to170.43%10
19usecarly.com170.43%16
20learn.microsoft.com160.40%15
21nextautomation.us160.40%10
22artisan.co150.38%12
23myaifrontdesk.com150.38%11
24salesmotion.io150.38%12
25support.microsoft.com150.38%13
26smartlead.ai140.35%11
27buildingradar.com130.33%5
28chromewebstore.google.com130.33%8
29get-alfred.ai130.33%7
30heyreach.io130.33%10

Why it matters

The standard advice that follows from the concentration story is to get onto the handful of domains that supposedly decide everything. For this class of question that advice is aimed at a target which does not exist. There is no gatekeeper list to buy your way onto: the long tail IS the market, most cited pages are cited once, and a large share of what gets cited is a vendor's own site. That is worse news for anyone selling a shortcut and better news for a company with something accurate to say about itself, because it means the entry cost is a page a retrieval system can actually use rather than a placement somebody else controls.

Method

  • A fixed universe of buyer questions was written in advance and frozen before any measurement ran. The questions are the kind a person types, not keyword strings.
  • Each question was put to two AI answer surfaces through their public APIs: OpenAI's search-grounded model and Claude with its web search tool.
  • Every response was stored whole, including responses where the model did not search the web and responses that failed. Failed and unsearched runs are counted in the denominator.
  • Every citation the surface reported was recorded with its host. Hosts were counted once per response, so a page cited three times in one answer counts once.
  • Requests that failed on our own rate limits were retried, and any that still failed were tagged and excluded from the rates while remaining visible as an instrumentation-failure count.

What this cannot tell you

Limitations

  • These are API surfaces, not the consumer products. The OpenAI search API is a close proxy for ChatGPT search and is not ChatGPT: the consumer product has memory, personalisation and its own system prompt. The same is true of Claude. Nothing here should be read as a measurement of a consumer app.
  • Generative answers are non-deterministic. Published work measuring this found day-to-day overlap of cited sources between 0.34 and 0.42, so a single observation of any one question means very little. Rates here are over many observations and are reported with their sample size.
  • The question universe is ours. It is weighted toward the categories DFX operates in, so this is not a map of all software buying. The concentration finding in particular may not hold for consumer questions, for medical or legal questions, or for any category where a small number of authoritative institutions genuinely exist. Published claims that Reddit and Wikipedia dominate AI citations are usually drawn from general-knowledge question sets, and that is a different measurement rather than a wrong one.
  • Two surfaces is not the market. Google's AI Overviews, Gemini, Perplexity and Copilot are not in this dataset, and published work finds surprisingly little overlap between surfaces.
  • This is a single point in time. It says nothing yet about trend.

The data

Published under CC BY 4.0. Use it, and cite DFX Intelligence with a link to this page.

A note on who published this

DFX Intelligence sells an AI operator, so we have an obvious interest in this subject. Two things follow from that and we would rather state them than have them noticed. First, the measurement includes our own product and reports its results without adjustment: at the time of publication we appear in none of the answers in this corpus. Second, the question set is ours and is weighted toward the categories we operate in, which is stated in the limitations rather than buried.

Report published 2026-08-19, updated 2026-08-19. Product facts referenced on this page were last verified 2026-08-19.

Facts on this page were last checked against the running product on . Prices are read from the live billing configuration, so this page and your invoice cannot disagree. The machine-readable version of this record is at /ai/entity.json.