Gemini and Claude visibility
By Olha Kryva ·
Gemini and Claude answer the same buying questions differently from each other and from ChatGPT, and a business visible in one is frequently invisible in another. There is no separate playbook for either; there is a reason to measure them separately instead of reporting one number that is wrong about most of them.
On this page 20 sections
- The short version
- How far apart are they, really?
- Why do they diverge?
- What this means for measurement
- What actually works for Gemini and Claude
- Find your weakest surface, then look at what it read
- Get the facts consistent
- Make your pages answerable
- Publish something checkable
- Gemini specifically
- Claude specifically
- A worked example
- Common mistakes
- Frequently asked
- How to read a four-assistant report
- What a low share on one assistant does not mean
- When the whole category is quiet
- Questions people actually ask
- Where to start
- One thing to do tomorrow
The short version
Gemini and Claude answer the same buying questions differently from each other and from ChatGPT, and a business visible in one is frequently invisible in another. There is no separate playbook for either.
What there is, is a reason to measure them separately rather than reporting a single “AI visibility” number that is wrong about most of them.
How far apart are they, really?
This is the part people underestimate, so here are the numbers rather than the adjective.
In one local market, four assistants were asked the same set of buying questions, five repeats each. Between them they named 63 businesses. Only four were named by all four. Forty-six appeared in a single assistant only.
In the same measurement, shares for one business by assistant:
| Assistant | Share of answers naming the business |
|---|---|
| Gemini | 100% |
| Perplexity | 83% |
| ChatGPT | 67% |
| Claude | 50% |
Identical questions, asked the same way, at the same time. An average across the four would have reported something near 75% — a number that described none of them and that nobody could act on.
Why do they diverge?
Partly observable, partly not. The honest split:
Retrieval differs. What each system fetches at question time, and from where. Perplexity retrieves heavily and cites; others lean more on what the model already holds.
Training and recency differ. A business that became prominent recently is more likely to appear in a system that retrieves than in one answering from memory.
Guardrails differ. Assistants vary in how willing they are to name specific businesses at all. In a category where one declines to recommend anyone, your share there will be structurally lower — and that is a property of the category, not a failing of your marketing.
Phrasing sensitivity differs. The same question rewritten slightly moves the answer more in some systems than others.
None of this is published, so the honest position is: we can observe the divergence reliably and explain it only partly. Anyone offering a confident causal account of why Claude names you less is guessing.
What this means for measurement
Report per assistant. Always. An average is a number that cannot be acted on — you cannot fix “the average”. The only useful version of the data keeps them separate.
Expect the ranking between assistants to change. Shares move independently. A business that is weakest on Claude this month may be weakest on ChatGPT next month, without anyone having done anything.
Do not chase one engine. A programme tuned to whichever assistant looked worst last month will spend its budget reversing itself.
Ignore small movement, especially on a single engine. Under about five points is drift, and the temptation to over-read one engine’s swing is strongest exactly when you have four numbers in front of you.
What actually works for Gemini and Claude
The same levers as everywhere else, because there is no separate playbook.
Find your weakest surface, then look at what it read
Where sources are shown, they are the shortest path to the work. Where they are not — and for these two they often are not — the pattern across answers still points at the source types carrying weight in the category. The businesses named instead of you are described somewhere, and that somewhere is usually a short list of directories and roundups.
Get the facts consistent
Assistants cross-check. A business described three different ways in three places gives a model less reason to name it than one described identically everywhere. This is the most reliable lever for systems whose retrieval you cannot observe, precisely because it affects every source at once.
Make your pages answerable
Buying question as the heading, answer in the first two sentences, plain words. Where the page is read at all, this is what determines whether it supplies the passage.
Publish something checkable
Specific numbers, stated method, named examples. This matters more for systems answering partly from training data, because being written about specifically is what puts you there in the first place.
The ordered version of the work is in how to improve AI visibility.
Gemini specifically
Two things worth knowing, both observable.
It is the one most tied to Google’s wider index. In practice, businesses with strong conventional search presence and well-maintained Google Business Profiles tend to do better here than their off-site coverage alone would predict.
It hedges less in local recommendation questions than Claude does, in the categories we have measured. That does not make it more accurate — it makes it more likely to produce a list at all, which raises everyone’s shares including your competitors’.
Claude specifically
It hedges more. In several measured categories Claude is the most likely of the four to answer a recommendation question without naming anyone, or to name fewer businesses. A lower share on Claude therefore does not necessarily mean you are described worse; it may mean the whole category produces shorter lists.
It is sensitive to how the question is framed. The same question asked as a personal situation rather than a request for a list tends to produce different names. This is a reason to include both phrasings in a question set rather than a lever to pull.
Read a Claude figure against the category, not against the other assistants. If nobody in your market is named often on Claude, your 50% may be the best number on the board.
A worked example
From a measurement on one local market, business anonymised.
Baseline. Named in 38% of general recommendation answers overall, but the per-assistant spread was 100 / 83 / 67 / 50 as above, and 0% of price answers on every assistant.
The finding that mattered. Not the spread between assistants, which is normal, but the fact that price questions were zero everywhere. A gap present on all four points at the sources rather than at any one system’s quirks.
What was done. Listings corrected and completed on the four sources cited for price questions; one page rewritten so the price question is answered in the first two sentences.
What moved. Price answers went from zero to roughly a third — first on Perplexity, then over following weeks on the others, unevenly. General recommendation moved four points and was reported as no change.
What was not claimed. That the work caused it, or that the per-assistant differences meant anything about the work. They were there before and after.
Common mistakes
Averaging the four. Produces a number describing none of them.
Treating a low Claude share as a problem to fix. Check the category first. If nobody is named often there, the number is about the category.
Optimising for whichever looked worst. They move independently; you will be chasing.
Reporting only the best one. The temptation is obvious and it destroys trust the first time the client checks another assistant themselves.
Assuming ChatGPT represents the others. It represents roughly a quarter of the picture.
Frequently asked
Which assistant matters most? Whichever your buyers use, which you will not know without asking them. In the absence of that, measure all four and weight by nothing — report them separately and let the pattern speak.
Why is my share so different between them? Different retrieval, different training, different guardrails. Partly explainable, not fully.
Should I have separate strategies per assistant? No. Same work, measured separately.
Can I be named by one and not the others forever? Yes, and it is common: 46 of the 63 businesses in our measurement appeared in exactly one assistant.
How to read a four-assistant report
A first measurement lands as four numbers, and the instinct is to look at the lowest one. Resist it. Read them in this order instead.
First, look at the spread. A business at 100 / 83 / 67 / 50 is in a normal range. A business at 60 / 55 / 58 / 52 is unusually even, which usually means its visibility comes from something all four read — typically one dominant directory. A business at 90 / 10 / 8 / 5 is carried by a single system, and that is fragile.
Second, look at the question types, not the totals. A business strong everywhere on general recommendation and absent everywhere on price has a source problem, not an assistant problem. A gap that appears on all four points at what is written about you; a gap on one points at that system’s retrieval.
Third, look at who else is named. If the same two competitors appear on all four and you appear on one, the difference is not subtle and it is not going to be fixed by clever wording.
Fourth, look at the absentees. Businesses you consider serious competitors that appear nowhere are informative: they tell you the category’s information supply is thin, and that becoming the source is available.
What a low share on one assistant does not mean
It does not mean that assistant dislikes you. These systems have no view about your business; they have retrieval and training data.
It does not mean the others are wrong. They are answering from different material.
It does not mean you should rewrite for that assistant. There is no assistant-specific rewriting that survives contact with the next model update.
It does mean the material that system draws on describes you less well — and since you cannot see that material directly for Gemini or Claude, the practical response is the general one: make what is written about you consistent and specific everywhere, and measure again.
When the whole category is quiet
Some categories produce short lists or no names at all, particularly where an assistant is cautious: medical advice, legal representation, financial products, anything involving a regulated qualification.
If your measurement shows every business in a market at low shares, the finding is about the category. Two things follow.
Do not buy a programme aimed at raising a number the category caps. A supplier reporting “we can get you from 12% to 40%” in a category where the highest observed share is 18% is selling something they cannot deliver.
Do consider the opportunity. A quiet category means the information supply is thin. Publishing the clear, factual, checkable material that does not yet exist is slower than fixing listings, and it is the one thing that compounds.
Questions people actually ask
Which of the two should I care about more? Whichever your buyers use. In the absence of that knowledge, neither: measure both and report them separately.
Gemini names me and Claude does not. What is wrong? Probably nothing about you. Check whether Claude names anyone in your category — if the lists are short everywhere, the difference is the category’s, not yours.
Does Google Business Profile affect Gemini? It is tied more closely to Google’s wider index than the others, so a well-maintained profile is more likely to matter there. We can observe the correlation; we cannot show the mechanism.
Can I be named by Claude without having a website? In principle, if what is written about you elsewhere is enough. In practice, rarely for a business, and it is not a strategy.
How many assistants should I measure? All the ones your buyers plausibly use. Four is what we run by default, because dropping one is how you find out six months later that it was the one that mattered.
Where to start
Measure all four, separately, before deciding anything. The free check covers them and emails you what came back, per assistant, with no call attached.
The method explains why they are never averaged. Related: Perplexity optimization, how to get recommended by ChatGPT, and what AI visibility is.
One thing to do tomorrow
Ask your three most valuable buying questions in Gemini and in Claude, five times each, and write down every business named. Twenty minutes will tell you whether the two agree about your market — and in our experience they will not, which is the whole point of measuring them separately rather than buying a single number from anybody.
Questions
- How different are they really?
- In one market, four assistants asked the same questions named 63 businesses between them. Only four were named by all four, and 46 appeared in a single assistant only.
- Why not report an average across assistants?
- Because you cannot act on an average. Shares move independently, and a middling number describes none of the engines it came from.
- Should I optimise for the weakest one?
- Look at what it read before answering, rather than tuning for the engine itself. A programme aimed at whichever engine looked worst last month tends to reverse itself next month.