Same Shortlist, Different Proof: Why AI Engines Disagree on Sources

Two Engines Can Name the Same Vendor and Trust Completely Different Evidence

A procurement lead at a 400-person software company in Atlanta asks two AI systems the same question: “What is the best customer-data platform for a B2B SaaS company using Snowflake and Salesforce?” ChatGPT names three vendors. Google AI Overviews names two of the same three. At first glance, the answers look aligned.

Then the buyer opens the sources.

ChatGPT supports one vendor with its integration documentation and a product page. Google points to a technical partner, a comparison article, and a discussion thread. Perplexity, asked moments later, produces a similar shortlist but cites yet another set of publishers.

The names converge. The proof fragments.

This is a practical problem for U.S. B2B SaaS teams. A company may earn a place in several AI-generated shortlists while the engines disagree about which pages make that place credible. A content plan built around “getting into AI” as if every engine used the same evidence can therefore produce misleading results.

The better question is not simply, “Are we visible?” It is: Which engine names us, for which buying need, using which evidence, and how stable is that pattern?

This guide explains what recent studies suggest, what they cannot prove, and how to build an evidence portfolio without pretending there is one universal LLM playbook.

AI Research Is Now Part of the B2B Buying Journey

AI-assisted research is no longer a fringe behavior. G2’s 2026 AI Search Insight Report says 71% of the B2B software buyers in its survey used an AI chatbot somewhere in their research process, while 51% said they started with one more often than Google. The report covers more than 1,000 buyers, but it remains one company’s survey and depends on self-reported behavior. Read the G2 report and methodology.

Gartner reported a related pattern from an August–September 2025 survey of 645 B2B buyers: 45% used generative AI during a recent purchase. Those buyers still used an average of seven information sources, and 69% turned to a sales representative to validate AI-generated information. That matters. AI does not necessarily replace the rest of the journey; it can decide which companies enter the journey. See Gartner’s published findings.

For a SaaS company, the commercial risk appears before attribution software can see it. If an assistant excludes the company from the initial shortlist, there may be no site visit, no campaign touch, and no lost-opportunity record. If it includes the company but describes pricing or fit incorrectly, the buyer may reject the option before speaking with sales.

That is why AI visibility must be measured as an observable part of market consideration—not as a new name for organic traffic.

The 32% Brand Overlap and 4% Source Overlap Finding

One small 2026 benchmark makes the fragmentation easy to see. DerivateX submitted 15 buyer-intent prompts to ChatGPT and Google AI Overviews and recorded 402 citations. For open-ended prompts, the two systems named overlapping vendors 32% of the time, but shared the same cited pages only 4% of the time. In forced head-to-head comparisons, source overlap increased, but only to 30%.

Those figures come from a single agency study with a small prompt set. They are not a universal law, and the company has a commercial interest in AI visibility work. Treat them as a dated observation, not a permanent benchmark. The valuable finding is the shape of the result: shortlist agreement can be materially higher than source agreement. Read the DerivateX benchmark.

Why can this happen?

An AI answer has at least two related jobs. It must decide which entities plausibly answer the buyer’s need, and it must find material that supports the claims it writes. The first job can draw on broad, repeated associations between a company and a category. The second may depend on the current retrieval system, query wording, response mode, geography, freshness, and availability of a page at that moment.

Two engines may know the same category leaders while searching different indexes, issuing different follow-up queries, or preferring different page types. Even one engine can change its sources when a prompt becomes more specific.

This is not evidence that citations reveal an engine’s complete reasoning. A visible source documents what the interface chose to show beside an answer. It does not expose training data, every retrieved page, or a private ranking formula. Still, source differences are actionable because they show which public evidence currently survives retrieval.

Each Engine Has a Different Evidence Diet

A separate 2026 benchmark from Gadex examined 150 answers across ChatGPT, Gemini, and Perplexity and logged 1,576 citations. Its published analysis reported that 83.3% of ChatGPT citations in the sample came from vendor-owned domains, while 78.3% of Perplexity citations came from comparison, list, or third-party evaluation pages. Only 21.6% of citations overall came from a domain that also ranked in Google’s organic top ten for the query.

Again, these are findings from one commercial vendor’s snapshot, collected on August 11, 2026. They may change with model, retrieval, and interface updates. They also should not be combined with other datasets as if the prompt sets were identical. See the Gadex B2B citation benchmark.

The useful implication is not “ChatGPT always trusts company websites” or “Perplexity never does.” Absolute claims like those collapse under a single counterexample. The practical implication is that a page-type mix that performs in one sampled engine may not transfer cleanly to another.

ChatGPT: make first-party facts complete and quotable

When an engine leans toward vendor-owned pages, weak product information becomes expensive. A polished homepage is not enough. The company needs stable pages that answer the questions a buyer actually asks: supported integrations, implementation ownership, security status, customer size, limits, migration requirements, pricing units, and the circumstances in which the product is not a fit.

The strongest page is often not the broad category page. It may be the integration guide that states exactly what synchronizes, how often, in which direction, and with what limitation. It may be a security page with a current date and clearly scoped claims. It may be a pricing page that uses full sentences rather than hiding every number inside an interactive calculator.

Google AI Overviews: prepare for query fan-out

Google’s official documentation says AI Overviews and AI Mode use query fan-out, issuing multiple related searches across subtopics and data sources. Google also says standard Search eligibility applies; there is no special schema type or secret AI file that guarantees inclusion. Read Google’s official guidance for AI features.

For a prompt about software for a healthcare company, fan-out may create separate retrieval tasks around HIPAA, implementation, integrations, company size, pricing, support, and migration. A company can be credible on the broad category and absent from the sub-question that decides the answer.

This rewards depth with structure. Each important claim needs a canonical, crawlable page. Related pages need clear internal links. Dates and ownership should be visible. Important facts should not depend on a user opening a modal or completing a form.

Perplexity: expect visible comparison and synthesis

Perplexity’s answer format places sources close to individual claims, making source quality unusually visible to the user. In the Gadex sample, third-party comparison and list pages dominated. That does not mean a SaaS company should chase indiscriminate placements. It means buyers and retrieval systems both benefit from pages that compare options with a disclosed method, current evidence, and real trade-offs.

A useful independent comparison defines the buyer, criteria, evidence date, and limits. It explains why one product fits a 50-person startup while another fits a regulated enterprise. A thin page that declares ten winners without showing how they were selected supplies little verifiable information.

Community Sources Matter, but Not in Every Category

Community discussions appear frequently in several AI citation studies, but the pattern is uneven.

Prefer analyzed 25 B2B software prompts and reported that Reddit was the most frequently cited source in its ChatGPT sample and the second-most cited source in its Google AI sample, behind YouTube. The sample is small, and Prefer sells services connected to AI search, so the result should be read as directional. See Prefer’s published analysis.

A broader EMGI study of 1,486 queries across 18 software categories reported that direct Reddit citations in Google AI Overviews increased from 2.5% for early research prompts to 20.1% for high-intent prompts. It also found large category differences. A source type that matters for CRM selection may contribute little in another software market. Read the EMGI analysis and limitations.

The wrong conclusion is to manufacture discussions. The right conclusion is to observe the real questions and language buyers use in public communities, answer openly when appropriate, and improve the company’s own material where those discussions expose missing information.

A durable strategy does not depend on controlling a community. It learns from communities while respecting their rules and preserving genuine authorship.

Pricing Pages Are a Special Failure Point

Pricing questions behave differently from broad category questions. “Best customer support software for a 100-person company” asks for judgment. “How much does Product X cost for 25 seats?” asks for a literal fact.

Yet SaaS pricing pages routinely make that fact difficult to retrieve. They hide billing units in tooltips, show a monthly figure without saying whether annual billing is required, omit minimum seat counts, or replace numbers with “contact sales.” Old announcements and partner pages then remain available with outdated figures.

Lyra published a monitoring example in which AI systems misquoted a vendor’s price in 38% of observed answers during a 30-day period. That is one vendor’s study, not a market-wide rate, and it should not be repeated as a universal probability. It nevertheless illustrates a failure any SaaS team can test directly. See Lyra’s pricing-page analysis.

A machine-readable pricing page should state, in ordinary sentences:

Structured data can clarify information that is already visible. It cannot rescue an ambiguous page, and it cannot guarantee that an engine will select the page.

Build an Evidence Portfolio, Not One “GEO Page”

The fragmented-source pattern changes the unit of work. The answer is not one enormous article engineered for every prompt. It is a small portfolio in which each page has a clear evidentiary job.

Layer 1: canonical company facts

Maintain one authoritative location for each fact a buyer may need to verify. Product scope, integrations, security, pricing, implementation, data handling, support, and limitations should be current, linked, and written in complete sentences.

Assign an owner and verification date. When a fact changes, update the canonical page first, then locate conflicting partner and publisher pages. Contradictory information forces an engine to choose between versions and forces a buyer to wonder which one is true.

Layer 2: decision-specific explanation

Create pages for constrained buying questions, not only broad category keywords. A useful page might explain whether the product fits a U.S. healthcare SaaS company with 80 employees, a six-month deadline, and a specific technology stack. It should name the assumptions and explain the trade-offs.

This is also where honest “not a fit” statements help. Exclusions reduce ambiguity. A product built for enterprise teams should say when a small company will pay for capacity it cannot use.

Layer 3: independently checkable evidence

Original research, technical partner documentation, disclosed comparisons, and dated observation articles can validate claims that a company cannot independently prove about itself.

At LLM Recommend, an observation article records the exact prompt, engine or mode, date, verbatim answer, and visible citations. It does not claim to reveal a model’s hidden logic. The value is reproducibility: another person can rerun the prompt and see whether the observation still holds.

Layer 4: engine-specific measurement

Run the same prompt panel separately in each engine that matters. Do not merge the results into one opaque “AI visibility score.” A company can be present in ChatGPT, absent from Google AI Overviews, and described inaccurately in Perplexity at the same time.

Record named companies, order when meaningful, description, visible citations, and errors. Preserve the full output. A cropped favorable sentence is not a measurement record.

A 60-Day Plan for One Engine and One Buying Need

A narrow test produces clearer evidence than an all-engine campaign.

Days 1–7: define and record

Choose one engine and one high-intent phrase used by real U.S. buyers. Build 15 to 20 prompt variations around the same job, changing one constraint at a time: company size, industry, integration, budget, timeline, or compliance requirement.

Run the panel and save the date, mode, full answer, named vendors, description, and citations. Score presence, factual accuracy, source type, and durability separately.

Days 8–20: repair factual weaknesses

Map every claim the answer would need to make. Identify the best current source for each claim. Fix missing, vague, contradictory, inaccessible, or outdated company pages. Update internal links so related evidence is discoverable.

Do not publish five new articles before correcting a wrong price or discontinued integration. Retrieval amplifies clear facts, including clear mistakes.

Days 21–40: fill evidence gaps

Create only the assets the baseline showed were missing. That may mean a detailed integration page, a pricing FAQ, a method-led comparison, partner documentation, original data, or a dated observation article on an appropriate owned or partner publication.

Every external contribution should stand on its own editorial value. Repeating the same promotional copy across unrelated domains creates volume without useful corroboration.

Days 41–60: test whether the change holds

Repeat the original panel weekly. Keep prompts stable. Compare answers and source sets. Mark model or interface changes so you do not mistake a system update for the effect of your work.

Measure initial movement around day 30. At day 60, ask whether accurate presence survived several observations. That is the point of LLM Recommend’s one-engine, one-keyword performance model: reduce ambiguity before expanding the scope. Teams can start with a live AI visibility audit to establish the baseline.

Measure Four Outcomes Separately

A useful report should preserve distinctions that dashboards often blur.

  1. 1. Recommendation presence: Did the engine name the company for the defined buying need?
  2. 2. Description accuracy: Were fit, capabilities, pricing, and limitations stated correctly?
  3. 3. Source support: Which pages were displayed, and did they actually support the nearby claim?
  4. 4. Commercial follow-through: Did answer-engine referrals, self-reported discovery, qualified calls, or opportunities change?

These outcomes can move independently. A company can gain a citation without entering the shortlist. It can enter the shortlist without owning a citation. It can receive a visit without attracting the right buyer. It can influence a sale without receiving a measurable click.

Report by engine and prompt family. Show the denominator. Preserve negative observations. A trustworthy report makes it possible to see what did not improve.

Our citation ownership guide provides a more detailed scorecard for separating mentions, links, visits, and pipeline. The AI agent buying guide explains how constrained prompts can rebuild a shortlist.

What the Current Evidence Cannot Tell Us

The research in this article has real limits.

Most engine-comparison studies are published by companies that sell AI visibility products or services. Prompt selection, geography, account state, interface, model, and collection date can all change the result. Several datasets contain fewer than 200 prompts. None reveals a proprietary ranking formula. Cross-study percentages cannot be combined cleanly because the methods differ.

The engines also change quickly. A source pattern observed in August 2026 may shift after a retrieval or interface update. Personalization and experimentation can cause two users to receive different answers at the same time.

Visible citations do not prove causation. They show which sources an interface displayed, not every source retrieved or every factor used to generate the answer. Recommendation order is not always a stable rank. A single screenshot is not proof of durable visibility.

The responsible response is not to ignore the data. It is to use it proportionally: form a hypothesis, run a controlled observation, make a targeted improvement, and repeat the same observation.

The Bottom Line

AI engines can agree on the shortlist and disagree on the proof. That is not a contradiction. It is the normal result of different retrieval systems, source preferences, query expansion, and answer formats working over a changing web.

For U.S. B2B SaaS companies, the practical strategy is straightforward:

There is no universal page that wins every AI answer. There is a disciplined way to make each important claim easy to find, verify, and describe correctly.

Start with one engine. Start with one buying need. Find the evidence gap. Fix it. Then observe whether the answer changes and whether the change holds.

Frequently Asked Questions

Why do ChatGPT and Google AI Overviews cite different sources?

They use different retrieval systems, indexes, query expansions, response modes, and selection methods. They can recognize the same vendors while choosing different pages to support the written answer.

Should a SaaS company create different content for every AI engine?

Not entirely. Maintain one accurate foundation of company facts, then use engine-specific observations to identify missing page types or evidence. Avoid duplicating nearly identical pages solely for different assistants.

Does structured data guarantee an AI citation?

No. Structured data can clarify visible facts and help systems interpret a page, but it does not guarantee inclusion, a citation, or a recommendation.

How should we measure visibility across AI engines?

Use a fixed prompt panel and track recommendation presence, description accuracy, displayed sources, and durability separately for each engine. Connect those observations to qualified commercial outcomes without claiming unsupported causation.

How often should LLM visibility be checked?

Weekly checks are useful during a focused test. Keep prompts stable, record the date and mode, and mark system changes. Evaluate initial movement around 30 days and durability through at least 60 days.

What should a B2B SaaS company fix first?

Correct inaccurate or inaccessible first-party facts first—especially pricing, integrations, security, implementation, and product fit. Then fill the specific external evidence gaps revealed by the baseline.