Article

Jul 26, 2026

AEO for B2B SaaS: Get Recommended in ChatGPT, Perplexity & Google AI

How AI engines actually pick the products they recommend: the Three-Layer Recommendation Stack, the measured citation levers, and the program to run, with every stat traced to its source.

AEO for B2B SaaS: Get Recommended in ChatGPT, Perplexity & Google AI

Somewhere right now, a buyer you'll never see is typing "best tool for [the exact problem your product solves]" into ChatGPT. They'll get a confident, five-product answer in eight seconds. Then they'll book demos with two of the five and never open Google at all.

Were you in the answer? That's the whole game now, and it stopped being a fringe behavior sometime last year. 51% of B2B software buyers now start product research with an AI chatbot more often than Google, per G2's survey of 1,076 buyers. Twelve months earlier it was 29%.

This guide is the deep version of how that game works and how to win it. Real mechanisms, sourced numbers, and the exact program we'd run on a Monday.

What is AEO for B2B SaaS?

Answer engine optimization (AEO) is the work of making your product the answer AI engines give when buyers ask for recommendations. It spans your own site (content engineered to be cited), your entity (how consistently the whole web describes you), and the third-party sources engines pull recommendations from. You'll also hear it called GEO. Same discipline, and the acronym argument is a distraction from harder questions.

One sentence to keep, because everything below falls out of it. Google ranks pages. AI recommends brands.

You already know the definition, so let's spend the word count where almost nobody does. What actually happens between a buyer's prompt and your product showing up, or not showing up, in the answer?

How does an AI engine actually decide what to recommend?

There's a pipeline between prompt and shortlist, and none of what follows is guesswork. It's been mapped from Google's own documentation and patents, Mike King's teardown of how AI Mode works, reverse-engineered ChatGPT internals, and the peer-reviewed GEO paper from KDD 2024. We compress it into three layers, because every brand that's missing from AI answers is failing at exactly one of them.


The Three-Layer Recommendation Stack: how an AI engine turns a prompt into a shortlist, and the three places brands fall out

Layer 1: Model Memory

Before the engine searches anything, it rewrites the buyer's prompt into dozens of its own smaller searches. Google calls this query fan-out. ChatGPT does the same thing, running 5 to 10+ search rounds behind a single answer. Here's the part that matters for you. Those searches are written from what the model already believes about your category.

So if the model doesn't associate your brand with your problem space, your brand never even makes it into the searches. Nobody looked you up and rejected you. You were invisible before the looking started. Olivier de Segonzac, who analyzed ChatGPT's search behavior, put it plainly. "Being unknown to the model means being invisible before the search even starts."

What builds Model Memory? Mostly, being talked about. Ahrefs studied 75,000 brands and found brand mentions across the web predict AI visibility about three times better than backlinks do. The signal SEO spent twenty years optimizing has been lapped by the one PR was building all along.

Layer 2: Retrieval

When the engine does search, it stops behaving like Google in a second way. It doesn't rank your page. It reads individual passages, scores each one on how directly it answers the question, and only then decides what to cite. Where the page ranked barely matters by that point. One 548,000-page analysis found ChatGPT cites only about 15% of the pages it retrieves. The rest get read and tossed.

Liam Dunne of Discovered Labs, whose breakdown of AEO for B2B SaaS is itself one of the most-cited sources on this topic, describes the consequence. "A page on a DR 15 domain with almost no backlinks that directly and clearly answers the question can get cited over a DR 70 site with thousands of backlinks." His team watched Google's AI Overviews do the same thing with a separate retrieval system that can cite a page ranking #15 while ignoring the #1 result.

If you're a Series A team staring up at incumbents, this is the best news in the discipline. The authority moat that made SEO a rich company's game guards the wrong gate now.

Layer 3: Consensus

The last layer decides whether you get named. Before an engine recommends you to a buyer, it runs what amounts to a fact-check. It compares what you say about yourself against what G2, Reddit, review sites, and the rest of the web say about you. Agreement reads as trustworthy. Conflict reads as risk, and the model handles risk by dropping you and recommending someone whose story holds together.

Kevin Indig, whose citation research with Semrush spans thousands of domains, explains the why. LLMs "lean very heavily on third parties to form their answers... they're trying to build consensus from many sources." Ethan Smith of Graphite drew the strategic conclusion on Lenny's Podcast. In an AI answer, "you need to get mentioned as many times as possible," because the model summarizes many sources instead of crowning one blue link.

That's the stack. Miss Layer 1 and you're never searched for. Miss Layer 2 and you're read but not cited. Miss Layer 3 and you're cited but never named. Every tactic worth doing serves one of the three, which makes it a useful filter for every pitch you'll hear in this category, including ours.

Why does this hit B2B SaaS harder than everyone else?

Because SaaS buying runs on exactly the question format AI answers best. "Best CRM for a 20-person agency." "[Incumbent] alternatives for mid-market." Buyers used to do that synthesis themselves across review sites and listicles. Now the model does it in one shot, before you know the account exists. In 6sense's study of 4,000+ B2B buyers, 94% used an LLM somewhere in the buying process.

And the buyers who arrive from those conversations behave differently. In Graphite's Webflow engagement, visitors referred by LLMs signed up at 24% against 4% for Google organic. Semrush ran the same question across 500+ topics and valued the average AI visitor at 4.4x a traditional organic visitor. Single-site numbers deserve their caveat, and the direction is not subtle.


AI-referred visitors convert at 24% vs 4% for Google organic in Graphite's Webflow case study, with Semrush and Ahrefs corroboration

Smith's explanation for the gap matches what we see on demo calls. The buyer shows up "so primed because you're having a conversation with multiple follow-ups... so when you're going somewhere, it's probably highly qualified." The conversation is the funnel. The click is just the receipt.

Meanwhile the old front door keeps narrowing. Pew watched 68,879 real Google searches and found users click a result on just 8% of visits when an AI summary appears, down from 15% without one. SparkToro's 2026 data puts 68% of Google searches at zero clicks. Fewer clicks to fight over, and the shortlist forming somewhere you can't see. That's the squeeze.

The Invisible Pipeline: what your dashboards don't show you

Three measured gaps separate what you can see from what's deciding your pipeline. We call the pattern the Invisible Pipeline, and it's why competent teams conclude AEO "isn't working" while a competitor quietly takes the category.


Three measured gaps: Reddit is 27% of what ChatGPT reads but 0.35% of visible citations; 62% of AI citations never name the brand; 90% of AEO conversions land as direct

Invisible retrieval. Discovered Labs intercepted ChatGPT's actual search traffic and analyzed 144,000 citations. Reddit made up 27% of everything ChatGPT read for buying answers, more than Bing News, yet showed up in just 0.35% of visible citations. Dunne's summary of his own data still stops me. "99% of Reddit's influence on what ChatGPT recommends is invisible." Judge Reddit by the citations you can see and you'll undervalue it by roughly 80x. Most SaaS marketing teams are doing precisely that.

Invisible influence. Indig and Semrush's ghost-citation study found 62% of AI citations link a page as a source without ever naming the brand in the answer. Your comparison post can feed the answer while a competitor gets the mention. His verdict, verbatim. "AI mentions are far more important than citations." Order matters too, because roughly 75% of the time a buyer takes an AI shortlist, they pick whichever brand appears first.

Invisible attribution. Graphite measured this at n8n. GA4 credited AI with 1% of conversions. The post-signup survey said 9%. And 90% of AI-driven conversions never clicked a citation at all, because the buyer read the answer, typed the brand into the address bar, and landed as "direct." When 90% of a channel's conversions file themselves under the wrong name, last-touch analytics will tell you to kill the exact thing filling your pipeline. The fix is embarrassingly low-tech, and we make every client ship it in week one. Ask "how did you hear about us?" on the signup or demo form, with an AI option.

What actually earns citations? The levers, with effect sizes

This is where the field stops being folklore, because someone benchmarked it. The GEO paper (Aggarwal et al., KDD 2024) tested nine content modifications across 10,000 queries and measured citation visibility in generative answers. The winners and the loser, in one chart.


Measured lift in AI citation visibility: quotations +27.8%, statistics +25.9%, fluency +25.1%, cited sources +24.9%, keyword stuffing below baseline

The fuller evidence table, with the rest of the measured levers.

Lever

Measured effect

Source

Add quotations

+27.8% citation visibility

GEO paper, KDD 2024

Add statistics

+25.9%

GEO paper

Cite named sources

+24.9%, and +115% for pages ranked fifth

GEO paper

Brand mentions across the web

Predict AI visibility ~3x better than backlinks

Ahrefs, 75K brands

Answers placed in the first third of the page

44% of ChatGPT citations come from there

Citation analysis, 548K pages

Human-written content

86% of Google results, 82% of LLM citations

Graphite, 31,493 keywords

Keyword stuffing

Below baseline

GEO paper

Notice what the top three levers have in common? Quotes, numbers, named sources. The exact things most SaaS content strips out in the name of brand voice. The page you're reading is built on those levers on purpose, and yes, that's recursive, and yes, it works, which is the point.

Two more findings from that table deserve a beat. Lower-ranked sites gain more from these optimizations than top-ranked ones, which stacks with the Layer 2 story. This channel structurally favors the challenger. And Graphite's human-vs-AI data lands on an uncomfortable truth for everyone currently flooding their blog with generated posts. Pure AI content gets cited less and ranks worse, by a margin so wide it's not worth debating. Indig's version of the rule is the one to tape to the wall. "The more your content resembles an actual study, the more likely it is to get cited."

Do ChatGPT, Perplexity, and Google source answers differently?

Measurably, yes, and the differences decide where your first quarter of effort should go. Profound analyzed 680 million citations across the major engines, the largest public comparison anyone's run, and the three have distinct personalities.

ChatGPT reads like a librarian. It favors authoritative knowledge bases and established media. Wikipedia alone accounts for 47.9% of its top-10 citation share. It also budgets hard, pulling 3 to 10 sources per answer, and Profound's conversation data holds a tactical gem. The first question in a conversation is about 2.5x more likely to trigger citations than the tenth. Your buyer's opening prompt is where the sourcing happens.

Perplexity reads like a forum lurker. It has the heaviest community bias of the three. Reddit is 46.7% of its top-10 citation share, roughly 3x ChatGPT's rate, and in Graphite's Webflow work YouTube showed up in 32% of Perplexity citations against 5% for ChatGPT. If your category's subreddits and YouTube reviews don't know you exist, neither does Perplexity.

Google's AI Overviews read like Google, minus the ranking loyalty. They pull from the regular Search index with the most evenly distributed sourcing of the three, but citation selection is decoupled from rank, which is how a #15 page gets quoted over the #1. Classic organic strength gets you into the pool. Passage-level clarity gets you cited. And as of June 2026, Search Console reports your generative-AI performance directly, so this engine is also the easiest to measure.

Our read on the sequencing, and we'll flag it as strategy inferred from the citation data rather than vendor-documented fact. Start where feedback is fastest. AI Overviews and Perplexity re-retrieve constantly, so content and community work show up there in days to weeks. ChatGPT's librarian bias leans on authority and mentions that accumulate slowly, so it pays out last and longest. Teams that invert this order spend six months waiting on the slowest engine while the two fast ones sit unworked.

What does an AEO program actually involve, stage by stage?

Five stages. Each produces an artifact you can inspect, and each serves a layer of the stack. If an agency can't show you the artifact, the stage doesn't exist.

Stage

Layer it serves

The artifact you see

Demand mapping

All three

Every buying query and prompt scored by buyer intent, reasoning visible per keyword

Answer-ready content

Retrieval

Pages engineered passage-by-passage to be cited, ranking in classic search too

Entity coherence

Memory + Consensus

An entity audit that finds every inconsistency across your site, profiles, and directories, with the fix list

Cited-surface placement

Memory + Consensus

The Citation Graph for your category, plus a log of placements won on it

Measurement

All three

Answer Share tracked across engines, tied to traffic and pipeline, every claim traced to a receipt

Two failure modes account for most disappointing engagements. Content without the offsite consensus work produces a well-written site nobody cites. Offsite mentions without measurement produce activity nobody can defend in a budget meeting. Ask any agency how the stages connect, and the answer tells you whether you're buying a system or a content calendar.

How do you measure a channel that hides its own influence?

Three practices, in the order you should ship them.

Survey attribution, this week. The n8n study validated it. Organic and paid proportions matched between GA4 and the survey, so the survey's AI number is trustworthy where last-touch is blind. One form field recovers 9x more visibility than your analytics has today.

Distribution logging, not spot checks. Engines are probabilistic, and they're personalized. ChatGPT builds a memory of who you are, which means checking your own visibility from your own account tells you nothing about what buyers see. Dunne's rule from the video is the operative one. "You want to log the distribution." Run your buying prompts across engines in volume, on clean sessions, and track Answer Share as a trend. Any single reading is noise.

Tie it to pipeline or don't bother. Answer Share is the leading indicator. Demos from buyers you never had to interrupt are the point. A report that can't trace a claim to a measurement is a vibe, and vibes are free elsewhere.

How fast does it move, and does it last?

Faster than SEO on one layer, slower on another, and the difference is the three-layer model. Retrieval moves quickly because there's no authority toll booth in front of passage selection. Discovered Labs reports citation movement in 24 to 72 hours for retrieval-layer fixes. Model Memory moves in quarters, because mentions accumulate slowly and models retrain on their own schedule. Set expectations by layer and nobody gets surprised in month two.

Decay is real on both ends. Answers refresh, competitors publish, citations age out. And there's a systemic wrinkle worth knowing about. Graphite found that roughly 40% of the citations ChatGPT already consumes are themselves AI-generated, and when models retrieve their own output, answer quality measurably collapses. The engines know this, which is why every citation study keeps finding the same preference for original, human, evidence-dense sources. The moat in this channel is having something real to say. Budget for the system and its upkeep, or don't start.

What should you do Monday morning?

Five moves, in order. Most are free.

  1. Run the cold-start test. Ask each major engine your category's ten most important buying questions from clean sessions. Where you don't appear, note which layer failed. Never retrieved is a Memory problem. Retrieved but not cited is a Retrieval problem. Cited but described wrong is a Consensus problem.

  2. Add the survey field. "How did you hear about us?" with an AI option, on every signup and demo form. You'll have your first real attribution data in two weeks.

  3. Pull the Citation Graph. For your top ten buying queries, list every source the AI answers cite. That list, not a keyword spreadsheet, is your placement target list for the quarter.

  4. Fix your three worst entity conflicts. Old brand names, stale pricing on review profiles, category descriptions that disagree with your homepage. Consensus is a machine reading the whole web, and it flunks contradictions.

  5. Rebuild one page around the levers. Take your most important comparison page and add real quotes, real numbers, and named sources, with the direct answer in the first third. That's a +25% class intervention, benchmarked, for an afternoon of work.

"This is all noise and none of it is proven"

The strongest objection to this whole field deserves a straight answer. Most of what's published about AEO is unvalidated, the engines are non-deterministic, and half the vendor studies collapse under one methodological question. Smith, whose agency runs control-and-test experiments across hundreds of queries before believing anything, says most practitioner claims in this space fail exactly that bar.

Which is why this page argues from the sources that survive it. A peer-reviewed benchmark. A 75,000-brand correlation study. Intercepted engine traffic. Real-behavior clickstream data from Pew. Case studies with disclosed methodology. Where a number is single-site, we said so. Where a claim is directional, we said that too. And two widely repeated conversion stats you've probably seen elsewhere failed source-tracing and sat out of this article entirely. The honest version of the objection kills the hype and leaves the mechanism standing. Buyers moved, engines decide by measurable rules, and the rules can be worked. As Smith put it when OpenAI's own team waved the "just make great content" flag, "anything can be optimized. You just need to understand the underlying systems and the rules of the game."

FAQ

Is AEO the same as GEO? Functionally yes. Same discipline, different acronym, and the vendors arguing about the name are usually avoiding harder questions. Evaluate the system and its artifacts.

How long does AEO take to work? Retrieval-layer changes can show citation movement in days. Model Memory takes quarters. Leading indicators typically move in the four-to-eight week range. Anyone quoting guaranteed day counts is manufacturing certainty the engines don't offer.

We're a small SaaS. Can we actually win against incumbents? Often more easily than you could in classic search. Citation doesn't care much about domain authority, mentions beat backlinks as the predictor, and the GEO benchmark found lower-ranked sites gain the most from optimization. Specific beats big on hundreds of buying questions, and most categories' AI answers are being settled right now.

Should we be on Reddit? Yes, and for a different reason than your social team thinks. Reddit is 27% of what ChatGPT reads for buying answers and holds formal data partnerships with both OpenAI and Google. Treat it as the training ground for your category's consensus, with genuinely useful participation. Mass-posted spam gets banned, and Smith notes the engines trust Reddit precisely because the community polices it.

Do AI visibility tracking tools work? They're directional, which is valuable if you treat them that way. Log distributions across many prompts, watch trends, ignore any single reading. The mistake is treating visibility as the goal instead of the leading indicator in front of traffic and pipeline.

How do we measure whether it's working? Answer Share across engines for your buying prompts. Citations of your pages in AI answers. AI referrals as their own analytics channel. Survey attribution to catch the 90% that arrives as direct. And, ultimately, pipeline.

Want the roadmap for your category? One call, and you leave with the map. The prompts worth winning, the pages to ship first, and how it all turns into pipeline. Get your roadmap.

Unlock your top growth channel

Book a call to start driving pipeline through AI search.