What 250 million AI search results say gets cited
Traditional SEO metrics explained only 4% to 7% of citation variance across 1,311 pages.
The 4 to 7 percent problem
Traditional SEO metrics still moved citations in the expected direction. Doubling those metrics was associated with roughly 25 to 40 percent more citations, and the relationship was statistically reliable at p<0.001. It was also weak. The models explained 4 to 7 percent of the variation in citation counts.
Traditional SEO metrics leave at least 93% unexplained
Share of citation variance explained across 1,311 pages.
- Variance explained by SEO metrics
- 4–7%
- Variance left to other factors
- 93–96%
Source: Profound analysis of 1,311 pages
That does not make links, rankings, or domain authority useless. It means they are table stakes with diminishing returns. A strong domain can improve the odds, but it cannot tell an answer engine which passage resolves the query, whether the claim is current, or whether the page looks useful from the retrieval snippet.
The relationship was real. It just left almost the entire citation decision unexplained.
The formats engines kept choosing
We classified 8,500 citations by page type. Blogs and opinion pages had moved ahead of comparative pages and listicles, but the useful grouping was the combination: those two formats supplied 61.5 percent of citations.
Blogs and comparisons supplied 61.5% of classified citations
Share of 8,500 citations by the type of page that received the citation.
- Blogs and opinion
- 34.2%
- Comparative pages and listicles
- 27.3%
- Documentation and wikis
- 15.8%
- Commercial and store pages
- 14.8%
- Community and forums
- 5%
- Homepages
- 2.2%
Source: Profound content classification, 8,500 citations
Both formats do a job that answer engines need. They name the question, gather candidate answers, and make a recommendation in a few self-contained passages. A homepage rarely does that. Homepages made up 2.2 percent of this sample.
Freshness mattered too. Fifty percent of top-cited pages were less than 13 weeks old. That gives teams a practical job: rerun the work, replace stale examples, and make the current answer easy to locate.
Retrieval happens before a model reads the page
During retrieval, an answer engine may decide between candidates from a title, description, URL, and a short snippet. In our tests, that snippet was around 100 characters. The full article cannot rescue a candidate that looks irrelevant at this stage.
We compared 50,000 highly cited URLs with 50,000 low-cited URLs. Slugs written as four to seven natural-language words were 11.4 percent more common in the highly cited group. URLs that were semantically closer to the query received up to 5 percent more citations.
The URL is a small retrieval document. So are the title and description. Each one should say what the page answers without making the retriever decode an internal ID or a vague brand phrase.
The same retrieval behavior helps explain the amount of user-generated content in AI answers. Forums and video transcripts contain plain-language questions, direct recommendations, and the exact phrases people use when they are deciding what to buy.
User-generated content share varied sharply by engine
Percentage of each platform's citations classified as user-generated content.
- ChatGPT17.4%
- Perplexity15.8%
- Google AI Overviews12.3%
- Microsoft Copilot4.6%
Source: Profound answer-engine citation analysis
One prompt becomes several searches
A ChatGPT prompt is not a keyword with extra words. The system turns it into query fanout, then searches those branches. In the observed sample, 36.4 percent of prompts produced two searches and 52.9 percent produced three. Nearly nine in ten produced two or three.
Most prompts generated two or three searches
Share of prompts by number of generated searches.
- Two queries
- 36.4%
- Three queries
- 52.9%
- One, four, or five queries
- 10.7%
Source: Profound ChatGPT query-fanout analysis
Across 1,000 SERP analyses and 1,000 ChatGPT executions, ChatGPT query fanouts overlapped 39 percent with Google results.
This is why optimizing only for the exact wording of a prompt misses the retrieval path. The better content brief asks which follow-up searches the model needs to answer before it can respond with confidence.
Commerce made the pattern harder to ignore
The product analysis compared 16,000 product detail pages queried from October 2 through November 2, 2025. The most frequently shown products had 848 percent more FAQs than the least frequently shown group. They had 103 percent more videos. Ratings, specifications, descriptive titles, availability signals, and readable URLs all moved in the same direction.
Frequently shown products supplied more answerable detail
Difference between the most and least frequently shown product pages.
- +848%
- FAQs
- +103%
- Videos
- Product rating36% higher
- Specification entries23% more
- Product title length18% longer
- Price11% higher
- Natural-language URLs7.7% more
- Discounts shown5.3% more
Source: Profound analysis of 16,000 product detail pages
The FAQ number is extreme, but the direction is ordinary. Product pages win when they expose the facts needed to finish a specific decision. A feed with current inventory, reviews, dimensions, and clear question-and-answer fields gives the agent less work to do.
The work I would prioritize
I would keep the SEO foundation, then spend the next unit of effort on retrieval clarity. Give the page a title and URL that name the question. Put the answer in a passage that can survive outside the page. Update it when the facts change. For products, expose the fields an agent needs to compare options without guessing.
Traditional SEO remains the foundation. The 4 to 7 percent result shows how much work begins after that foundation. Traditional SEO metrics left at least 93 percent of citation variance unexplained.
Study notes
- Dataset
- More than 250 million responses and 3 billion citations from frontend monitoring across eight answer engines.
- Published subsamples: 1,311 pages; 8,500 citations; 50,000 top-cited and 50,000 bottom-cited URLs; 1,000 Google SERP analyses and 1,000 ChatGPT executions; and 16,000 product detail pages.
- Collection window
- Product detail pages were queried from October 2 to November 2, 2025.
- Engines and products
- ChatGPT, Perplexity, Google AI Overviews, Google AI Mode, Google Gemini, Microsoft Copilot, Claude, Meta AI
- Sample
- The published query analyses used commercial and informational intent.
- The URL study compared top-cited with bottom-cited pages; the commerce study compared the most with the least frequently shown product pages.
- Analysis
- Statistical association analysis of traditional SEO metrics and citation counts, reporting explained variance and a p-value.
- Content classification, high-versus-low citation URL comparison, ChatGPT fanout and Google SERP overlap, and descriptive product-page comparison.
- Contributors
- Josh Blyskal, author and presenter
- Access
- Limitations
- The article combines analyses with different units and subsamples, so its percentages are not estimates from one common sample.
- The 4–7% figure describes measured associations across 1,311 pages; it does not establish a causal effect of SEO metrics on citations.
- Details not published
- A common collection window, full engine-level counts, sampling frame, model specification, and raw-data access are not provided in the available deck or recording page.
Original research
- 1.
- 2.
Questions about the research? Email Josh.