Data reference · Published August 3, 2026

AI search statistics and research findings

A citable reference to Josh Blyskal's original research on how ChatGPT, Claude, Google AI products, Perplexity, and other answer engines search, retrieve, cite, and represent information.

Every row preserves the reported number, unit, sample, date, source, and material limitation. These studies use overlapping corpora and different units; their sample sizes should not be added together.

By · AI Strategy & Research at Profound · Study-level credits below

Headline findings

8 citable answers

These are the findings most likely to answer a direct question. The answer text in this table also appears in the page's FAQPage structured data.

Swipe horizontally to see the study, sample, and source for each answer.

Eight headline findings from Josh Blyskal's AI search research
QuestionAnswerFindingStudySampleSource and link
What is the largest category of ChatGPT prompt intent?37.5%of classified prompts asked ChatGPT to create, draft, or complete something.50M+ ChatGPT prompt studyClassified sample from 50M+ prompts; classified n unpublishedSourcePermalink
How much AI citation variance did traditional SEO metrics explain?4–7%of citation variance was explained by the tested traditional SEO metrics.250M-response analysis1,311 pagesSourcePermalink
Which page formats supplied most classified AI citations?61.5%of classified citations came from blogs, opinion, comparisons, or listicles.250M-response analysis8,500 classified citationsSourcePermalink
How did frequently shown product pages differ from the least shown group?+848% FAQson frequently shown product pages than on the least shown group.250M-response analysis16,000 product-detail pagesSourcePermalink
What share of Claude citations also appeared in Brave's top 10?79.2%of Claude's cited URLs appeared in Brave's first ten results.State of AEO 2026~35,000 URLs across 400 queriesSourcePermalink
How much did Claude and ChatGPT citation domains overlap?8%average citation-domain overlap between Claude and ChatGPT.State of AEO 2026600+ queries; domain-level comparisonSourcePermalink
What was the most-cited domain across the tracked answer engines?3.11%aggregate citation share made Reddit the most-cited domain in the measured set.Reddit citation study4B+ citations and 300M responsesSourcePermalink
Were natural-language URL slugs more common among highly cited pages?+11.4%more common among highly cited URLs for four-to-seven-word slugs.250M-response analysis50,000 highly cited vs. 50,000 low-cited URLsSourcePermalink

Reading standard

How to read these numbers

Treat every finding as a statement about its measured sample, not a universal law of AI search.

Answer engines change by model, mode, index, prompt, location, and date. A percentage from Claude cannot be transferred to ChatGPT. A result about cited domains does not automatically describe cited URLs, answer language, clicks, or purchases.

The most reliable rows have a public denominator and a defined unit. When a source publishes a percentage but not the subgroup size, this page says so. When an analysis is observational, the wording uses associated with or more common rather than caused.

The compendium favors primary work authored or co-authored by Josh. Public archive claims without enough information to identify the measure or sample were reviewed but left out. A smaller set of inspectable numbers is more useful than a long list of unsupported statistics.

Metric definitions

Web-search invocation
The share of observed responses in which the product called a live web-search tool.
Citation share
The fraction of observed citations attributed to a domain or content class, not the share of all answers.
Domain overlap
The share of domains common to two measured source sets. It does not require matching URLs or passages.
Query fanout
The underlying searches an answer engine generates from one user prompt before composing an answer.
Intent share
The fraction of classified prompts assigned to an intent category, not the fraction of unique users.
Referral traffic
Visits carrying an answer-engine referrer. Citations and unlinked mentions can influence users without creating one.

What 50 million ChatGPT prompts reveal about user intent

June 25, 2025

A prompt-intent study separating generative, informational, commercial, transactional, navigational, and connective conversational turns.

More than 50 million real ChatGPT prompts surfaced through Profound Prompt Volumes. The published percentages describe a classified sample drawn from that corpus.

Method

Prompt-intent classification, category-share calculation, and comparison of four shared categories with a published traditional-search baseline.

By Josh Blyskal and Sartaj Rajpal

Swipe horizontally to see samples, dates, method notes, and sources.

What 50 million ChatGPT prompts reveal about user intent findings
MeasureResultFindingSample and dateMethod noteSource and link
Generative intent37.5%Generative requests were the largest category, five percentage points above informational prompts.

Classified sample from a 50M+ prompt corpus; classified n unpublished

Collection window unpublished; report published June 25, 2025

A generative prompt asks the model to produce an output, not simply retrieve information.Full studyPermalink
Informational intent32.7%Almost one-third of classified prompts asked for information or explanation.

Same classified sample; n unpublished

Collection window unpublished

The share was 20 percentage points below the traditional-search baseline used.Full studyPermalink
Navigational intent2.1%Navigational prompts were 30.1 percentage points below the 32.2% traditional-search baseline.

Same classified sample; n unpublished

Collection window unpublished

The baseline source and its sampling method are not identified in the public report.Full studyPermalink
Transactional intent6.1%Transactional prompts were 5.5 percentage points above the 0.6% traditional-search baseline.

Same classified sample; n unpublished

Collection window unpublished

A prompt can express purchase intent even when the transaction happens elsewhere.Full studyPermalink
Connective turns12.1%Turns such as thanks, please, and revision requests did not fit a conventional search intent.

Same classified sample; n unpublished

Collection window unpublished

Conversation logs contain dependent turns that keyword datasets usually omit.Full studyPermalink

The classified-sample size, collection window, labeling procedure, traditional-search baseline source, and raw data are not public.

The unit is prompts, not unique users. Follow-up turns are included, so the shares are not equivalent to first-turn search demand.

What 250 million AI search results say gets cited

December 8, 2025

A collection of analyses on SEO metrics, page formats, retrieval snippets, query fanout, freshness, URLs, and product-detail pages.

More than 250 million responses and 3 billion citations observed across ChatGPT, Perplexity, Google AI Overviews, Google AI Mode, Gemini, Copilot, Claude, and Meta AI.

Method

Several published subsamples were analyzed separately: association modeling, content classification, high-versus-low citation URL comparison, query-fanout comparison, and descriptive product-page analysis.

By Josh Blyskal

Swipe horizontally to see samples, dates, method notes, and sources.

What 250 million AI search results say gets cited findings
MeasureResultFindingSample and dateMethod noteSource and link
SEO metrics and citation variance4–7%The tested traditional SEO metrics explained a small share of variation in citation counts.

1,311 pages

Published December 8, 2025; analysis window unpublished

p<0.001, but association does not establish causation.Full studyPermalink
Doubling tested SEO metrics+25–40%Doubling the measured SEO metrics was associated with roughly 25% to 40% more citations.

1,311 pages

Analysis window unpublished

This is the same observational model as the 4–7% explained-variance result.Full studyPermalink
Blogs, opinion, comparisons, and listicles61.5%Blogs and opinion supplied 34.2%; comparisons and listicles supplied another 27.3%.

8,500 classified citations

Collection window unpublished

This is citation composition, not the probability that a page of each type will be cited.Full studyPermalink
Homepage citation share2.2%Homepages were the smallest named page-type category in the classified set.

8,500 classified citations

Collection window unpublished

The result does not include later changes in hyperlinking or referral behavior.Full studyPermalink
Age of top-cited pages50% <13 weeksHalf of the top-cited pages were less than thirteen weeks old.

Top-cited-page slice; subgroup denominator unpublished

Collection window unpublished

Page age is descriptive and may reflect query freshness or publication mix.Full studyPermalink
Four-to-seven-word URL slugs+11.4%Natural-language slugs of four to seven words were more common in the highly cited group.

50,000 highly cited URLs vs. 50,000 low-cited URLs

Collection window unpublished

Group prevalence does not prove that changing a URL causes citation growth.Full studyPermalink
URL similarity to the queryUp to +5%URLs semantically closer to the query received up to 5% more citations.

100,000-URL high-versus-low comparison

Collection window unpublished

The public materials do not publish the similarity model or confidence interval.Full studyPermalink
Prompts generating two or three searches89.3%36.4% generated two searches and 52.9% generated three.

Published ChatGPT fanout sample; n unpublished

Collection window unpublished

The result describes the measured prompt mix, not every ChatGPT mode or query.Full studyPermalink
ChatGPT fanout overlap with Google39%Generated ChatGPT search strings overlapped 39% with Google result sets.

1,000 Google SERP analyses and 1,000 ChatGPT executions

Collection window unpublished

The overlap unit and matching method should be read from the full study context.Full studyPermalink
Frequently shown product pages+848% FAQsThe most frequently shown product pages had 848% more FAQs and 103% more videos than the least shown group.

16,000 product-detail pages

October 2–November 2, 2025

The comparison is descriptive; it does not isolate the effect of FAQs or videos.Full studyPermalink

The findings come from different units and subsamples; they must not be treated as one pooled sample.

The 4–7% result is an observational association, not a causal estimate.

A common collection window, full engine-level counts, model specification, and raw data are not public.

The state of AEO in 2026: Claude is not ChatGPT

July 22, 2026

A set of engine-level analyses covering search invocation, source overlap, citation formats, query fanout, referral traffic, and AI search advertising.

Published analyses span Claude, ChatGPT, Brave Search, Google Search, and Google AI Mode. Each slice uses a different unit.

Method

Observed search invocation, matched Claude citations to Brave positions, compared domain overlap and content types, tracked repeated fanout strings, and separately observed referrals and ads.

By Josh Blyskal · Jasman Singh, research lead

Swipe horizontally to see samples, dates, method notes, and sources.

The state of AEO in 2026: Claude is not ChatGPT findings
MeasureResultFindingSample and dateMethod noteSource and link
Claude web-search invocation36.6%Claude searched for a little over one-third of tested prompts with search enabled.

Mixed recommendation and explainer prompt set; n unpublished

Collection date unpublished; report published July 22, 2026

The search rate depends on the tested prompt mix and should not be generalized universally.Full studyPermalink
Claude citations in Brave positions 1–1079.2%Nearly four in five cited URLs appeared on Brave's first results page.

Approximately 35,000 URLs across 400 queries

Collection date unpublished

The match does not by itself prove that Brave directly caused source selection.Full studyPermalink
Citation-domain overlap with Google's top 5064% vs. 37%Claude citation domains overlapped 64% with Google's top 50; ChatGPT citation domains overlapped 37%.

Pairwise domain comparison; n unpublished

Collection date unpublished

Against Google's top 10, the reported figures were 34% for Claude and 21% for ChatGPT.Full studyPermalink
Claude and ChatGPT citation-domain overlap8%The two answer engines shared very few cited domains on average.

More than 600 queries; pairwise domain comparison

Collection date unpublished

The overlap formula, distribution around the average, and per-engine citation counts are unpublished.Full studyPermalink
Listicle citation share36.4% vs. 19.7%Claude used more listicles than ChatGPT in the classified citation set.

Classified citations; n unpublished

Collection date unpublished

Source-type labels describe composition, not page-level citation probability.Full studyPermalink
Forum and UGC citation share0.9% vs. 15.8%Claude used far less forum and user-generated content than ChatGPT.

Classified citations; n unpublished

Collection date unpublished

The comparison is engine- and sample-specific.Full studyPermalink
Fanouts containing a year94% vs. 17%Claude added 2025 or 2026 to far more fanouts than ChatGPT.

Query-fanout sample; n unpublished

Collection date unpublished

The measured years and prompt mix make this a time-bound result.Full studyPermalink
Repeated Claude fanout strings~65%The same query strings recurred in roughly two-thirds of repeated Claude fanouts.

Repeated fanout executions; n unpublished

Collection date unpublished

The number of repetitions and matching rule are not public.Full studyPermalink
Observed ChatGPT referral traffic+60%Referral traffic rose roughly 60% overnight and settled near 1.6 times the prior global level.

Observed referral dataset; exact size unpublished

May 7–May 22, 2026 comparison

Referral behavior is a separate dataset from the retrieval and citation analyses.Full studyPermalink

Several slices still omit exact prompt or citation counts, retrieval-analysis dates, sampling procedures, and raw data.

The 36.6% search rate reflects the tested mix of current recommendations and basic explainers.

Prompt routing, citation overlap, content classification, referral traffic, and ad observations are separate analyses.

Why Reddit became AI search's most-cited domain

Nov 2025 · May 2026 update

An aggregate and engine-level analysis of Reddit citation share, source pairing, sentiment, community concentration, and cited-post age.

More than 4 billion citations and 300 million answer-engine responses in the main study, plus a separate follow-up using approximately 7 million recent ChatGPT citations and fanouts.

Method

Aggregate and engine-level domain ranking, source-pair review, sentiment-rate comparison, subreddit concentration, post-age analysis, and a later fanout trend comparison.

By Josh Blyskal and Sartaj Rajpal · Profound, in collaboration with Reddit

Swipe horizontally to see samples, dates, method notes, and sources.

Why Reddit became AI search's most-cited domain findings
MeasureResultFindingSample and dateMethod noteSource and link
Aggregate Reddit citation share3.11%Reddit ranked first among cited domains in the pooled engine dataset.

4B+ citations and 300M answer-engine responses

August 2024–late October 2025

No single domain supplied most citations; 3.11% led a fragmented source market.Full studyPermalink
Reddit rank by engineTop 3 on 5 of 6Reddit ranked first on Perplexity, second on ChatGPT, AI Overviews, and Grok, and third on AI Mode.

Six tracked answer engines

August 2024–late October 2025 aggregate window

Microsoft Copilot was the outlier, ranking Reddit number 31.Full studyPermalink
Brand-sentiment citation rate6.1% negative / 5.0% positiveNegative and positive brand commentary was cited at similar rates.

Reddit content containing brand sentiment; subgroup n unpublished

Observed in the 2025 study

The result does not measure the sentiment of all Reddit content or resulting answers.Full studyPermalink
Communities used per query class3–5 subredditsAnswer engines often concentrated retrieval within a few topic-specific communities.

Query-class analysis; n unpublished

Observed in the 2025 study

The named communities varied by topic and purchase context.Full studyPermalink
Average age of a cited Reddit post~1 yearFour percent of cited posts were published in 2019 or earlier.

Cited Reddit posts observed in 2025; n unpublished

Post-age findings observed in 2025

ChatGPT's cited set peaked in Q1 2025; Perplexity's peaked in Q1 2024.Full studyPermalink
ChatGPT fanouts explicitly adding Reddit0.15% → 3.68%The share rose approximately twenty-fourfold between January and late May.

ChatGPT fanout trend plus ~7M recent citations

January–late May 2026

Country, language, industry, and prompt composition were not published.LinkedInPermalink
Reddit share of recent ChatGPT citations8.5%Reddit returned to the number-one cited domain in the separate follow-up.

Approximately 7M recent ChatGPT citations

Published June 2, 2026

This is a later ChatGPT-only slice, not the 3.11% six-engine aggregate.LinkedInPermalink

The aggregate result pools answer engines with materially different behavior.

Exact engine-level counts, subgroup sizes, labels, and raw data are not public.

Citation share measures sourcing visibility, not whether a user saw, trusted, clicked, or acted on an answer.

The May 2026 fanout finding is a separate ChatGPT-only follow-up and should not be pooled with the 2025 six-engine analysis.

SAGE: the operating method behind the research

Published July 26, 2026

SAGE is a four-stage method for turning answer-engine observations into repeatable work: Setup, Analyze, Generate, and Engineer.

It is a practitioner method, not an empirical dataset. The numbers below are operating heuristics or linked research evidence and should not be presented as a controlled validation of the framework.

Provenance

Invented by Josh Blyskal at Profound and taught in Profound 101. The method keeps teams from automating an AEO workflow before its prompts, diagnosis, and handoff make sense by hand.

Swipe horizontally to see each stage's question, process, and output.

The four stages of the SAGE method for AEO
StageQuestionWhat happensOutput
01SetupCould I explain why every topic and prompt belongs in this setup?Setup is where I decide what the measurement is actually for. A large prompt list does not reassure me, because I have seen plenty of large setups that nobody on the team has read. I start with the short category terms customers use, check the demand around them, read real prompt examples, and then write a small set that covers the category and the buyer questions we care about. I should be able to explain why each prompt is there before I add another hundred.A baseline with a clear reason for every topic, prompt, competitor, and filter.
02AnalyzeWhat changed, where did it change, and which source helps explain it?When a dashboard changes, I start with the prompts underneath the number. I read the answers, separate the engines, and inspect the exact pages they cited. Sometimes the explanation is a new publisher, sometimes it is a competitor page, and sometimes the engine simply changed how it framed the question. The analysis is finished when I can describe what happened in plain language and point to the evidence behind it.A prioritized set of gaps with evidence behind each diagnosis.
03GenerateWhat can we publish or change that responds to the gap we found?Once the gap is specific enough to explain, I decide what kind of work would address it. A new page is one option, although an existing page often needs a clearer answer or a narrower job. There are also topics where publishers, communities, or product information supply most of the evidence, which means another article on the brand site may do very little. The diagnosis should determine the work, including where that work lives.A published page or another concrete change that responds to the diagnosed gap.
04EngineerWhich part of the process has been repeated enough that the team can trust it?I leave automation until the team has run the workflow by hand more than once and agrees on what a good result looks like. Automating earlier makes the same unresolved judgment call recur at a higher speed. Once the sequence is familiar, I automate the repetitive collection and reporting while leaving the diagnosis and decision with a person.A repeatable workflow that keeps a person responsible for the decision.

Evidence used inside SAGE

Selected field studies from the archive

July 2025–July 2026

These public analyses have clear numeric claims and identifiable samples but do not yet have full study pages on this site. They are included because they answer recurring practitioner questions about measurement, retrieval, citation speed, source diversity, and volatility.

Evidence tier

Treat these as public field notes. Each row links to the original post and states what the post did not disclose. They are not pooled with the four full studies above.

Swipe horizontally to see samples, dates, method notes, and original posts.

Selected AI search field studies from Josh Blyskal's archive
MeasureResultFindingSample and dateMethod noteSource and link
One run vs. ten runs per prompt per day10.24% vs. 9.99%In Profound economist Jennifer Zou's study, citation share differed by 0.25 percentage points despite ten times as many daily executions.

753 prompts, seven platforms, ~989,000 runs, and 6.66M citation slots

United States, June 1–14, 2026

The pooled portfolio result does not establish that one daily run is sufficient for an individual prompt.Profound studyPermalink
Daily movement after more prompt runs0.36 pp → 0.21 ppZou's portfolio analysis found that ten runs per prompt reduced daily citation-share movement by about 40%, but the remaining movement was already small.

Same 753-prompt, seven-platform, fourteen-day portfolio

United States, June 1–14, 2026

Portfolio averages can hide prompt-level and between-engine instability. The study also resampled 2,000 synthetic portfolios.Profound studyPermalink
Time to first ChatGPT or Claude citation6.81-day medianThe 75th percentile was 18.68 days and the 90th percentile was 37.10 days.

Approximately 900 newly published marketing pages

Published May 11, 2026; observation window unpublished

The public post does not explain uncited-page handling, prompt exposure, engine splits, page selection, indexing lag, or collection method.LinkedInPermalink
ChatGPT web-search invocation by intent53.5% commercialCommercial prompts searched at 53.5%, informational prompts at 18.7%, and generative prompts at 8.9%.

667,000 ChatGPT conversations

Published January 8, 2026; collection window unpublished

The public post does not disclose language, geography, industry mix, intent-labeling method, or first-turn handling.LinkedInPermalink
Overall ChatGPT web-search invocation17.4%ChatGPT used live web search in fewer than one in five conversations in the dataset.

667,000 ChatGPT conversations

Published January 8, 2026; collection window unpublished

The result describes this conversation mix and model period, not all ChatGPT usage.LinkedInPermalink
Unique cited domains per prompt8.98 vs. 5.12Google AI Mode cited more unique domains per prompt than ChatGPT in the comparison.

19M Google AI Mode citations; ChatGPT comparison sample size unpublished

Published July 9, 2025; collection window unpublished

The public post does not disclose prompt mix, geography, or whether both products used matched prompts.LinkedInPermalink
July cited domains absent in June40.5–59.3%The share was 59.3% for Google AI Overviews, 54.1% for ChatGPT, 53.4% for Copilot, and 40.5% for Perplexity.

Approximately 80,000 prompts; platform-level denominator ambiguous

June 11–13 vs. July 11–13, 2025

Sources differ on whether ~80,000 prompts were tested per platform or across all four; this is an asymmetric new-in-July metric.Profound studyPermalink

Methodology

Source map

This page is a structured index of published findings, not a new pooled analysis. The underlying studies used different collection systems, samples, periods, engines, and units. The underlying raw data is not distributed or licensed through this page.

Swipe horizontally to compare every study's unit, sample disclosure, and method.

Methodology and sample disclosure for every research group
StudyPublishedUnitSample disclosureMethod
50M+ ChatGPT prompt studyJune 25, 2025Classified promptsSample drawn from 50M+ prompts; classified n and collection window not publishedIntent classification and category-share comparison
250M-response analysisDecember 8, 2025Pages, citations, URLs, prompts, SERPs, and product pages250M+ responses and 3B citations overall; published subsamples range from 1,311 pages to 100,000 URLsAssociation, classification, comparison, and descriptive analyses
State of AEO 2026July 22, 2026Prompts, cited URLs, domains, fanouts, and referralsSome slices report ~35,000 URLs across 400 queries or 600+ queries; other exact counts remain unpublishedRouting observation, pairwise overlap, classification, and trend analysis
Reddit citation studyNovember 10, 2025Responses, citations, domains, posts, and query classes4B+ citations and 300M responses; subgroup counts not publishedRanking, source pairing, sentiment, concentration, and post-age analysis
SAGE methodJuly 26, 2026Operating stagesPractitioner method; no empirical sampleSetup, Analyze, Generate, and Engineer operating loop
Selected field studiesJuly 2025–July 2026Prompts, responses, citations, conversations, pages, and domainsReported separately in each row; no pooled samplePortfolio comparison, time-to-event summaries, routing rates, and trend comparisons

Inclusion standard

  • Josh authored, co-authored, presented, or publicly documented the analysis.
  • The claim has a numeric value and an identifiable unit.
  • A primary public source remains accessible.
  • Missing denominators, dates, or procedures are stated rather than inferred.

What this page does not claim

  • The samples are not representative of every AI answer or user.
  • Observed associations do not prove that a tactic caused the result.
  • Citation share is not equivalent to user attention, traffic, trust, or revenue.
  • The public materials do not provide enough raw data or procedural detail for independent replication.
  • Engine behavior measured on one date may change after a product update.

How to cite this research

Citation guide

Cite this compendium when referencing the collection. Cite the individual study when using a specific finding so readers can inspect its methods and limitations.

APA-style

Blyskal, J. (2026, August 3). AI search statistics and research findings. JoshBlyskal.com. https://www.joshblyskal.com/research/findings

Suggested in-text citation: (Blyskal, 2026). For a direct statistic, include the study name and measured sample in the surrounding sentence.

BibTeX

@misc{blyskal2026aisearchfindings,
  author       = {Josh Blyskal},
  title        = {AI Search Statistics and Research Findings},
  year         = {2026},
  month        = {August},
  howpublished = {\url{https://www.joshblyskal.com/research/findings}},
  note         = {Published August 3, 2026. Accessed YYYY-MM-DD}
}

Individual study citations

  1. Blyskal, J., & Rajpal, S. (2025, June 25). What 50 million ChatGPT prompts reveal about user intent. JoshBlyskal.com.
  2. Blyskal, J. (2025, December 8). What 250 million AI search results say gets cited. JoshBlyskal.com.
  3. Blyskal, J. (2026, July 22). The state of AEO in 2026: Claude is not ChatGPT. JoshBlyskal.com. Research lead: Jasman Singh.
  4. Blyskal, J., & Rajpal, S. (2025, November 10). Why Reddit became AI search's most-cited domain. JoshBlyskal.com.
  5. Blyskal, J. (2026, July 26). SAGE for AEO: A Four-Stage Operating Loop. JoshBlyskal.com.

Do not strip the sample from the statistic.

Write “4% to 7% across 1,311 pages,” not simply “SEO explains 7% of AI citations.” The unit and denominator are part of the finding.