Home Blog Insights Case Study Round-Up: What the Data Actually Shows About SaaS Companies Winning in AI Search
Insights

Case Study Round-Up: What the Data Actually Shows About SaaS Companies Winning in AI Search

We pulled the six published datasets with real samples and stated methods (Ahrefs, two Siege Media studies, Semrush, Similarweb, and a 973-site ecommerce paper) to separate what AI search actually does for B2B SaaS from the anonymous visibility-score screenshots. Here is what each one proves, where it stops, and how to become your own case study in 90 days.

Alex Carter
Alex Carter
August 20, 2026
14 min read
3,137 words
Case Study Round-Up: What the Data Actually Shows About SaaS Companies Winning in AI Search

Six published datasets carry real, verifiable numbers on how companies actually perform in AI search: Ahrefs, two separate Siege Media studies, Semrush, Similarweb, and a working paper covering 973 ecommerce sites. Together they show AI referral traffic is small, growing fast, and wildly inconsistent in how it converts. Most vendor case studies you will find cite none of this.

Why almost every “AI search case study” is unusable

Search for proof that generative engine optimization works and you get a wall of agency landing pages claiming a client went from 8% to 24% visibility, or 0 to 23.4%, or 16 to 74%. Almost none of them name the client. Almost none define the metric. Almost none say how many prompts were tested, on which models, in which week, or whether the “before” measurement was taken on a Tuesday when ChatGPT happened to be routing to a different retrieval path.

That is not a small methodological quibble. AI visibility scores are enormously sensitive to prompt set, model version, personalization, and whether the search tool fired at all. Semrush’s clickstream data found ChatGPT enabled its search feature on just 34.5% of queries as of February 2026, down from 46% in late 2024. A visibility score measured across prompts where search never fired is measuring the model’s training data, not your optimization work. Two agencies can run the same audit on the same site in the same month and produce numbers that differ by 40 points, and both can be reporting honestly.

So we went looking for the opposite: published research with a stated sample, a stated window, a stated definition, and numbers someone would have to retract if they were wrong. Six qualified. Here is what each one actually proves, where it stops, and what a B2B SaaS team should copy from it.

If you want a defensible baseline for your own site before you read further, our GEO audit runs a fixed buyer-intent prompt set across the major engines so your “before” number is reproducible. No obligation to do anything with it.

The six datasets that hold up

1. Ahrefs: 0.5% of traffic produced 12.1% of signups

Ahrefs published its own numbers, which is rare and worth more than a hundred anonymized client wins. Over a 30-day window, 0.5% of Ahrefs traffic came from AI search while 12.1% of Ahrefs signups did. That works out to AI search visitors converting at roughly 23 times the rate of traditional organic search visitors, for Ahrefs specifically.

The supporting behavioral data is the part most summaries skip. AI search visitors had a lower bounce rate and viewed 50% more pages per visit, but spent less time on site per visit. And users in AI search clicked links 75% less often than they did in traditional organic search. The majority of both the traffic and the growth came from ChatGPT.

What it proves: for a high-consideration B2B SaaS product with a well-known brand and a free signup, AI-referred visitors can arrive dramatically further along the buying journey. They have already read a synthesized comparison. They are not browsing.

What it does not prove: that your company will see 23x. This is one company, one month, one product category, and that category is SEO tools, which is the single most over-documented topic on the internet and therefore the easiest for a model to answer confidently. Ahrefs is also the brand most likely to be named in an SEO tool comparison regardless of any optimization. Treat 23x as a ceiling for an unusually favorable case, not a benchmark.

What to copy: the measurement setup. Ahrefs could report this because it separated AI referrals as their own channel and tied them to signups, not to sessions. Most teams cannot answer this question at all because AI referrals are landing in “direct” or getting lumped into “referral” in GA4. Fixing attribution is a prerequisite, not a nice-to-have.

2. Siege Media: comparison pages predict AI traffic better than anything else

This is the most actionable study of the six. Siege Media pulled 116 B2B GA4 properties and correlated the volume of transactional comparison content each site had built against AI-referred sessions over the trailing 90 days ending May 24, 2026. Across those properties, AI search referred 739,492 total sessions, spread across 1,112 transactional pages.

Comparison pages were defined concretely: URLs containing “vs,” “alternative,” or “best…software,” with a minimum of 50 sessions in the window, after removing false positives and deduplicating. AI search meant chatgpt.com, chat.openai.com, openai.com, perplexity.ai, claude.ai, gemini.google.com, copilot.microsoft.com, or meta.ai.

The result, segmented by how many such pages a site had:

Comparison pages on site Median AI search sessions
1 to 5 Baseline
6 to 20 350% increase over baseline
21 or more 900% total increase over baseline

What it proves: content type is a stronger lever than most of what gets sold as GEO. Sites with real depth in “vs” and “alternatives” and “best X software” content pull meaningfully more AI-referred sessions than sites without it.

What it does not prove: causation, and the study is careful about this. Sites with 21+ comparison pages are probably also bigger, better funded, more linked, and more likely to be named in the underlying training data. Publishing 21 thin comparison pages next quarter will not produce a 900% lift.

What to copy: the strategic read is that bottom-of-funnel comparison content is now doing double duty. These are the same pages that already convert well in classic SEO, so the investment case does not depend on AI search working out. That is the rare GEO tactic with no downside if the whole channel underdelivers. We cover the mechanics of building these in our guide on how to get cited by ChatGPT and Perplexity.

Worth pausing here: if your team has fewer than five comparison pages live, that is likely the highest-leverage gap on your site right now, and it is a content problem rather than a technical one. If you want a second opinion on which comparisons to build first, a short working session will get you a prioritized list faster than a tool will.

3. Siege Media again: homepages are absorbing the traffic

A separate Siege Media study looked at 50 national websites, 28 B2B and 22 B2C, comparing the three months ending May 19, 2025 against the prior three months in Search Console. Homepage impressions rose 10.7%, and homepage clicks rose even as sitewide traffic declined. B2B showed stronger homepage growth than B2C.

The study states plainly that causality is impossible to prove in aggregate, which is exactly the disclosure the anonymous agency case studies never make. But the pattern is consistent across brands and lines up with how AI Overviews and chat assistants behave: when a model is asked about a company rather than a topic, it points at the front door.

What it proves: the homepage is now an answer surface, not a brand poster. If a model has to describe what your company does and who it is for, and your homepage says “Unlock your potential,” the model has nothing to work with and will paraphrase a G2 listing instead.

What it does not prove: that homepage optimization causes AI traffic. Correlation over a period when many things changed at once.

What to copy: put the plain-language answer to “what is this and who is it for” in crawlable text above the fold, with the category name spelled out the way a buyer would say it. This is the cheapest fix in the entire GEO stack and most SaaS homepages fail it.

4. Semrush: 17 months of clickstream, and a concentration problem

Semrush analyzed more than one billion lines of US clickstream data from its 200 million user panel, covering October 2024 through February 2026. The headline is growth: outbound referral traffic from ChatGPT to the rest of the web grew 206% between January 2025 and January 2026.

The uncomfortable finding sits underneath it. Over 30% of all ChatGPT referral traffic goes to just 10 domains, and 21.6% of it goes to Google. So roughly a fifth of what gets counted as “AI referral traffic” is ChatGPT handing users back to a search engine.

The reach numbers tell a similar story. ChatGPT referred traffic to about 71,000 unique domains monthly in October 2024, peaked at 260,000 in October 2025, then settled back to 170,000 by February 2026. The long tail expanded and then contracted.

Two more findings change how you should think about keywords. Between 65% and 85% of ChatGPT prompts did not match any traditional search keyword in Semrush’s 27-billion-keyword database. And average queries per session rose 50% over the study’s final four months, reaching 1.75 by February 2026.

What it proves: the channel is growing fast and is highly concentrated. It also proves that keyword research as currently practiced misses most of the demand, because people phrase prompts conversationally and at a level of specificity that never showed up in search volume tools.

What it does not prove: that the concentration is permanent. It is a young distribution.

What to copy: stop building your prompt test set from your keyword list. Build it from actual sales call language, support tickets, and the questions prospects ask on demos. Those are the phrasings that match how people talk to models. If the vocabulary here is still fuzzy, our marketing glossary defines the terms without the acronym soup.

5. Similarweb: the growth rate is real and the base is tiny

Similarweb tracked referral traffic to the top 1,000 websites globally and estimated that AI platforms generated over 1.13 billion referral visits in June 2025, up 357% from June 2024. In the same month, Google search generated 191 billion. ChatGPT accounted for over 80% of all AI-driven referrals. Referrals specifically to news and media sites were up 770% year over year.

Hold those two numbers next to each other. AI referrals grew at a rate that would be extraordinary in any channel, and they still represent well under 1% of what Google sends. Both facts are true and neither one cancels the other.

What it proves: the trajectory justifies investment. The current volume does not justify reallocating your entire organic budget.

What it does not prove: anything about B2B SaaS specifically. This is top-1,000 global sites, which skews heavily toward news, retail, and reference. Your category’s curve may look nothing like it.

What to copy: the framing for your leadership team. “Growing 357% year over year off a base under 1%” is an honest sentence that survives scrutiny in a board meeting. “AI search is replacing Google” does not, and will cost you credibility in six months.

6. The ecommerce working paper: the counterexample nobody quotes

This is the most important entry on the list, because it points the other way. A working paper titled “ChatGPT Referrals to E-Commerce Websites: Do LLMs Outperform Traditional Channels?” analyzed 973 ecommerce websites generating $20 billion in combined revenue over 12 months, August 2024 through July 2025. Search Engine Land’s write-up reports that ChatGPT referral traffic accounted for roughly 0.2% of total sessions, about 200 times smaller than Google organic, and that affiliate and organic search both converted better than ChatGPT: affiliate by 86% and organic search by 13%.

The paper’s models projected continued gains for ChatGPT but no parity with organic search within the following year.

What it proves: the “AI traffic converts better” claim is not a law of nature. In ecommerce, at scale, across a year, it converted worse.

What it does not prove: that B2B SaaS behaves the same way. Ecommerce and B2B software have almost nothing in common in buying journey length, consideration depth, or what a model can usefully synthesize. A model comparing four project management platforms is doing genuine work for the buyer. A model recommending a phone case is not.

What to copy: the skepticism. If a vendor cites the 23x number without mentioning that a larger, longer study found the opposite in a different vertical, they are selling, not analyzing.

Where the evidence contradicts itself

Put Ahrefs and the ecommerce paper side by side and you get a 23x conversion advantage in one and a conversion disadvantage in the other. Both are real measurements. The reconciliation is not that one is wrong, it is that conversion advantage from AI referral is a function of how much pre-purchase synthesis the buyer needed.

Where the buying decision requires comparing feature sets, pricing tiers, integration support, and vendor reputation, the model does the qualifying work that used to take three blog posts and a demo. The visitor who clicks through has been filtered. Where the decision is fast and low-consideration, the model adds a step instead of removing one, and the referral looks like ordinary browsing traffic with extra friction.

The practical implication for B2B SaaS is that you are probably in the favorable case, but you should verify that with your own data rather than assuming it. Which brings us to the useful part.

What all six agree on

Source Sample Window Core finding
Ahrefs Own site, first-party 30 days 0.5% of traffic drove 12.1% of signups
Siege Media (versus pages) 116 B2B GA4 properties, 739,492 AI sessions 90 days to May 24, 2026 21+ comparison pages correlate with 900% higher median AI sessions
Siege Media (homepages) 50 sites, 28 B2B and 22 B2C 3 months to May 19, 2025 Homepage impressions up 10.7% while sitewide traffic fell
Semrush 1B+ clickstream lines, 200M panel Oct 2024 to Feb 2026 ChatGPT outbound referrals up 206%; 30% goes to 10 domains
Similarweb Top 1,000 global websites June 2024 to June 2025 1.13B AI referral visits, up 357%, versus 191B from Google
Ecommerce working paper 973 sites, $20B revenue Aug 2024 to Jul 2025 ChatGPT was 0.2% of sessions and converted below organic

Strip out the disagreements and four things survive every one of these studies.

ChatGPT is the channel. Every source that broke out platforms found ChatGPT dominant, at over 80% of AI referrals in the Similarweb data and the majority of both traffic and growth in the Ahrefs data. Optimizing for six engines equally is a misallocation.

Volume is still small in absolute terms. Nobody credible reports AI referral traffic as a large share of sessions. The numbers land under 1% in the ecommerce study and at 0.5% in the Ahrefs data.

Growth rates are genuinely steep. 206% in Semrush’s clickstream, 357% in Similarweb’s referral estimates. These are not rounding errors.

Measurement is the bottleneck. Every study that produced a useful number did so because someone had already separated AI referrals as a distinct channel. Teams that have not done this cannot participate in the conversation, let alone optimize.

That last point is where most engagements actually start. Our AI SEO practice spends the first two weeks on attribution before touching content, because a lift you cannot measure is a lift you cannot defend at renewal. You can also see how that has played out on real engagements.

How to become your own case study in 90 days

The honest conclusion from reviewing all six is that external benchmarks tell you the shape of the channel, not your position in it. Here is the sequence that produces a number you can actually stand behind.

Weeks 1 and 2: fix attribution. Create a channel group in GA4 that isolates referrals from chatgpt.com, chat.openai.com, perplexity.ai, claude.ai, gemini.google.com, copilot.microsoft.com, and meta.ai. Use the same list Siege Media used so your numbers are comparable to published research. Tie it to signups or qualified leads, not sessions. If you only measure sessions, you will conclude the channel is irrelevant, because by session volume it currently is.

Weeks 3 and 4: build a fixed prompt set. Forty to sixty buyer-intent prompts, written from sales call transcripts and support tickets rather than your keyword list, since Semrush found most prompts do not match traditional keywords at all. Freeze the set. Record the model and the date every time you run it. A visibility number without a frozen prompt set is not a measurement. Our GEO audit checklist walks through the run protocol.

Weeks 5 through 10: build comparison depth. This is the Siege Media finding applied. Audit how many real “vs,” “alternatives to,” and “best X software” pages you have live. If the answer is under five, that is your quarter. Build them properly, with actual feature comparisons and honest treatment of where competitors are better, because models reward specificity and readers punish obvious spin. These pages earn their keep in classic organic search regardless of what happens in AI search.

Weeks 5 through 10, in parallel: fix the homepage. Plain-language category description in crawlable text, above the fold. Who it is for. What it replaces. Pricing signal if you have one. This takes a day and it is the highest ratio of impact to effort on the list.

Weeks 11 and 12: re-run and compare. Same prompt set, same models, log the date. Report the delta alongside the AI-referred signup number from your GA4 channel group. Two independent measurements that move together are a finding. One that moves alone is noise.

Ninety days is realistic for a first honest reading, not for a transformation. Anyone promising a visibility transformation in 30 days is measuring something that moves on its own. If you want the plan pressure-tested against your specific category before you commit a quarter to it, book a call and we will go through it with you.

What to ignore

A short list, based on what did not survive the filter for this article.

Unnamed client percentages. “We took a client from 8% to 24% visibility.” Which client, which prompts, which model, which week. Without those four, the number is decoration.

Single-run visibility scores. Models are non-deterministic and personalize. One run is an anecdote. Three runs across three days on a frozen prompt set is a measurement.

Cross-industry conversion benchmarks. The gap between Ahrefs at 23x and the ecommerce paper below parity is the whole argument. Any single multiplier presented as a universal benchmark is ignoring one of those two.

Anything measuring the wrong engines. If a report weighs six engines equally when one of them carries over 80% of referrals, the composite score is diluted to the point of uselessness.

Anything that treats GEO as separate from SEO. The strongest predictor in the strongest study was comparison content, which is a bottom-of-funnel SEO asset that has existed for fifteen years. If you want the longer version of that argument, start with our full guide to generative engine optimization.

The short version

The companies genuinely winning in AI search are not the ones with the best visibility score screenshot. They are the ones who separated AI referrals as a channel early, built real comparison depth because it paid off in organic search anyway, wrote a homepage a model can quote, and measured on a frozen prompt set they did not change every time the number moved in an inconvenient direction.

None of that is exotic. All of it is verifiable. And when the next wave of case studies arrives with numbers attached, you will have your own data to check them against, which is worth considerably more than trusting theirs.

If you would rather not build the measurement apparatus yourself, our GEO audit produces the frozen prompt set, the baseline, and the comparison-content gap analysis in one pass. Take the findings and run with them in-house if that suits you better.

Alex Carter
Alex Carter LinkedIn
SEO & Content Strategy, MV3 Marketing

Alex Carter leads SEO and content strategy at MV3 Marketing, specializing in generative engine optimization, technical SEO, and AI-driven content systems for B2B companies.

Ready to audit your organic growth opportunity?

$2,500 flat. 5 business days. Six deliverables tied to pipeline , not rankings. No retainer required.

Get the Organic Growth Audit →

Turn Your Organic Channel into a Revenue Engine.

The MV3 SEO Audit maps your full organic opportunity in 5 business days.

Get the Audit →