The question usually arrives in a particular shape. A founder has read that Perplexity takes something close to half its citations from Reddit, that G2 and Capterra are the gatekeepers of a software category, and that the fix is schema markup and an FAQ block. Three of those four things are either wrong or unsupported by the studies they are drawn from.
The short answer
Almost nothing Perplexity quotes for a category question is a page the vendor controls. In a study of roughly 350,000 articles across 52 B2B SaaS categories, run in May 2026 against ChatGPT, Claude, Perplexity and Google AI Overviews, earned media accounted for 61 percent of AI citations and brand-owned content for 29 percent. 1
That single split decides the work. If the pages being quoted are mostly written by other people, the job is not publishing more of your own. It is being present, by name and accurately, in the pages other people already publish, and being the version of the answer that is easiest to lift. 1
Which sources actually appear
The clearest breakdown by source type comes from a smaller and narrower test: 100 SaaS shopping keywords pulled through DataForSEO on 6 May 2026, US English desktop, taking the first three citations from each answer. It measures Google AI Overviews rather than Perplexity, so read it as the shape of the category rather than as Perplexity's own numbers. 2
| Source type | Citations | Share |
|---|---|---|
| Vendor listicle or blog, published by a third party | 150 | 51% |
| Forum and community, including Reddit and Quora | 43 | 15% |
| Vendor product or category page | 33 | 11% |
| YouTube | 28 | 10% |
| Analyst, including Gartner, Forrester and IDC | 16 | 5% |
| Review aggregator | 9 | 3% |
| News and media | 9 | 3% |
| Other | 6 | 2% |
The first line is the one to sit with. Half of everything cited was a listicle or blog post published by somebody other than the vendor being asked about. That is the surface the original question points at, and it leads the table by a wide margin. 2
The review site assumption does not hold
Across those same 100 queries, every well-known review aggregator combined, meaning G2, Capterra, TrustRadius, Software Advice, GetApp, PeerSpot and Business.com together, earned nine citations. 2 For a channel that dominates how B2B buyers are told to research software, that is a small footprint.
The Reddit number that gets repeated is not the one in the study
The figure in circulation puts Reddit somewhere between a fifth and a half of Perplexity's citations. The primary study says something much narrower. Across 217,000 unique prompts and 248,000 unique cited Reddit URLs, refreshed in October 2025, Reddit is the single most cited domain on Perplexity and appears in only about 3.5 percent of Perplexity answers. 3
Both statements hold at once. Reddit leads the domain table because citations spread thinly across a very long tail, not because it saturates the answers. The same body of work puts Perplexity's most cited domains as Reddit, LinkedIn, NIH, Microsoft and Google. 4 When Reddit does appear on Perplexity it lands early, at an average position of 3.4, against 6.7 on SearchGPT and 8.8 on Google AI Mode. 3
The threads that get cited are old and quiet
This is the finding that contradicts most of the Reddit work sold alongside AI visibility retainers. The posts being quoted are not the ones that did well.
- 80 percent of cited posts had fewer than 20 upvotes, and 70 percent had fewer than 20 comments. 3
- Median engagement on a cited post was 5 to 8 upvotes and 11 to 19 comments. 3
- The average cited post was around 900 days old. 3
- Question and answer threads were more than half of all Reddit citations, and Q&A, comparison and discussion formats together came to close to three quarters. 3
A thread from two and a half years ago carrying six upvotes is not a campaign outcome. It is a durable answer to a question people keep asking, which is a different asset and is built differently. Chasing velocity optimises for the thing the data says is not being selected.
What the cited pages have in common
The 350,000-article study compared cited pages against uncited ones on the same keywords. The gaps that showed up are unglamorous. 1
| Property | Cited | Not cited |
|---|---|---|
| Statistics per article | 4.2 | 1.2 |
| Expert quotes per article | 1.6 | 0.2 |
| Word count | 1,690 | 1,305 |
| Reading grade level | 9.6 | 10.8 |
Denser in checkable numbers, carrying named human quotes, moderately longer, and easier to read. None of that is a trick. It describes a page an engine can lift a sentence from without having to interpret it first.
What did not correlate
Two of the most commonly sold fixes did not separate cited pages from uncited ones in that study. Schema markup showed no gradient across rankings at all. FAQ sections correlated weakly, appearing on 24 percent of cited articles against 17 percent of uncited ones. 1
The gate that comes before any of it
Perplexity runs two agents and they behave differently. PerplexityBot is described in Perplexity's own documentation as designed to surface and link websites in search results on Perplexity, is stated as not used to crawl content for foundation models, and obeys robots.txt. Perplexity-User fetches a page when somebody's question needs it, and generally ignores robots.txt because a person initiated the request. Perplexity's own recommendation is to allow PerplexityBot if you want to appear in its results. 5
Allowing it guarantees nothing. It is a precondition. A blocked crawler means the index behind cited answers never sees the page, and no amount of work on the page changes that.
The order this actually goes in
- 01Confirm PerplexityBot is allowed in robots.txt and not blocked at the firewall. It is a five-minute check and everything else depends on it. 5
- 02Find the third-party listicles already ranking for your category questions. Half of what gets cited for shopping queries is that page type, published by somebody who is not you. 2
- 03Get named in them accurately, and get the outdated entries corrected. An entry describing a product you shipped two years ago is worse than no entry.
- 04Answer the durable questions in the communities where they keep being asked, and accept that the useful thread may be years old and quiet. 3
- 05On your own pages, raise the density of checkable figures and named quotes rather than the word count. The gap between cited and uncited was 4.2 statistics against 1.2. 1
- 06Keep the review profiles current, and do not treat them as the route in. Nine citations across 100 shopping queries is the size of that channel. 2
How to tell whether it worked
The measurement problem is real and worth stating plainly. Answers are generated fresh, so the same prompt can return different sources on different days, and a single check proves nothing in either direction. Any honest measurement runs a fixed prompt set repeatedly, from a baseline taken before the work starts, and reads the sources rather than the sentiment. A screenshot of one good answer is not evidence.
Receipts
05 sources- 01Citera, 350,000 B2B SaaS articlesRoughly 350,000 articles drawn from 10,382 keywords across 52 B2B SaaS categories, tested against ChatGPT, Claude, Perplexity and Google AI Overviews, US English, May 2026. Source of the 61 percent earned media against 29 percent owned split, the 19 against 9 percent review platform figure, the 4.2 against 1.2 statistics and 1.6 against 0.2 expert quote gaps, the 1,690 against 1,305 word counts, the 9.6 against 10.8 reading grades, the absence of a schema gradient and the 24 against 17 percent FAQ figure.Last checked 3 September 2026
- 02AUQ.io, 100 SaaS shopping queriesFirst three citations from Google AI Overviews for 100 SaaS shopping keywords, pulled via DataForSEO on 6 May 2026, US English desktop, 294 citations in total. Source of the source-type table and of the nine combined citations earned by every well-known review aggregator. Measures Google AI Overviews rather than Perplexity, which is why it is read here as the shape of the category.Last checked 3 September 2026
- 03Semrush, 248,000 cited Reddit URLs217,000 unique prompts across Google AI Mode, Perplexity and ChatGPT Search, identifying 248,000 unique cited Reddit URLs, data refreshed October 2025 and published 10 November 2025. Source of the 3.5 percent of Perplexity answers figure, the 3.4 average citation position against 6.7 and 8.8, the 80 percent under 20 upvotes and 70 percent under 20 comments, the 5 to 8 and 11 to 19 medians, the roughly 900 day average post age and the Q&A share of citations.Last checked 3 September 2026
- 04Semrush, most cited domains in AI230,000 prompts across ChatGPT Search, Google AI Mode and Perplexity over the 13 weeks from 14 July to 12 October 2025, tracking more than 100 million citations, published 10 November 2025. Source of Perplexity's most cited domains being Reddit, LinkedIn, NIH, Microsoft and Google.Last checked 3 September 2026
- 05Perplexity crawler documentationPerplexity's own developer documentation on its two agents. PerplexityBot is described as designed to surface and link websites in search results on Perplexity, is stated as not used to crawl content for AI foundation models, and obeys robots.txt, with the recommendation to allow it so that a site appears in results. Perplexity-User supports user actions and generally ignores robots.txt because a person initiated the request.Last checked 3 September 2026

