AI Search Visibility Study: Who Wins Citations in ChatGPT, Perplexity & Google AI
How 3,600 SaaS-related prompts across ChatGPT, Perplexity, Claude, and Google AI Overviews cite sources — and what predicts inclusion.
Executive summary
Problem — SaaS teams don't know what drives citations inside AI answer engines. Traditional SERP ranking has been decoupled from LLM inclusion — and no operator has a reproducible model for it.
Why it matters — For SaaS, AI answer engines already intercept high-intent research queries. Being un-cited is the modern equivalent of being un-indexed.
- SEO and content leads adapting to generative engines
- Founders evaluating GEO consultancies
- PR teams measuring earned mentions
- Median SaaS domain earns 0.4 citations per 100 prompts — a floor most teams don't measure.
- First-party data pages get cited 3.8× more than opinion posts.
- Structured data + author bylines are the single largest citation predictor after topical authority.
0.4 for the median SaaS site; top 10% earn 6.2.
Pages with original data get 3.8× more citations than opinion pieces.
Named-author pages with LinkedIn schema saw 2.4× the inclusion rate.
Perplexity cites articles ≥2,400 words 71% of the time.
83% of ChatGPT citations are ≤18 months old.
Median answer cites 4.1 unique domains — up from 2.8 a year ago.
Pages with FAQ + Article + Author schema showed a 1.9× citation lift.
One 'category-leader' domain earns 34% of all citations within its vertical.
Research objectives
- Quantify citation share for SaaS content in AI answer engines.
- Identify the strongest predictors of inclusion.
- Track drift across ChatGPT, Perplexity, Claude, and Google AI Overviews.
- What content attributes predict citation inclusion?
- How much do engines differ in their citation biases?
- Does traditional SEO ranking still predict AI citation?
Methodology
- 3,600 systematically generated SaaS prompts across 12 buyer intents
- Citation extraction from ChatGPT, Perplexity, Claude, Google AI Overviews
- Enrichment via Ahrefs, Common Crawl, and public schema data
- Prompts stratified across informational, transactional, and comparison intent
- Excluded prompts triggering safety refusals
- Logistic regression to identify citation predictors
- Shapley value analysis for feature importance
- Chi-square tests for engine-level bias
- Replicated prompt sample 3× at 30-day intervals
- Independent citation coding by 2 analysts (κ = 0.87)
- Prompt bias — the panel reflects buyer-journey prompts, not casual queries.
- Engine drift within the 6-month window.
Data & visualizations
Analysis & insights
- 01Being 'cited by AI' is a category-share game. If you don't own a category, you don't get cited in it.
- 02Original research is the highest-leverage GEO investment. Everything else is compounding on that base.
- 03Author entity clarity is under-priced — most SaaS sites still ship pages with no author schema.
Recommendations
The 3.8× citation lift comes from measured, sourced data — not opinion.
Combined 1.9× lift is the cheapest GEO win available.
ChatGPT freshness bias filters out 83% of stale content.
Concentration means one dominant page can capture 34% of vertical citations.
- Engines change ranking behavior without notice — results may drift within 90 days.
- US-en only in this cut; localization studies pending.
Download the full research package
FAQ
What is GEO?
Generative Engine Optimization — the practice of optimizing content for inclusion in LLM-powered answer engines.
Does traditional SEO still matter?
Yes. Traditional SERP ranking correlates positively with AI citation, but no longer perfectly.
How do you extract citations?
Automated extraction with human review at 10% sample for accuracy validation (κ = 0.87).
Get the monthly research digest
One email per month with the latest SaaS marketing research, benchmarks, and experiments. No spam.
Join the newsletter