Method and limits

Is Sayso accurate?

Accurate about what it measures, within limits we publish instead of hiding. This page sets out exactly what the number is, how it is produced, where the error comes from, how big each source of it is, and what we do about each one. Nothing here is a selling point we would drop in a bad quarter.

The number

A share of answers, per assistant

Everything we report is the same simple fraction: how often an assistant named your business out of the number of times we asked.

  • What it is

    Times named ÷ times asked, calculated separately for each assistant and each type of question. A count you could do by hand, which is the point: you can check it.

  • What it is not

    Not a rank — there is no position inside a written answer. Not a score out of a hundred. Not an average across assistants: four numbers that disagree do not become one number that is right.

  • Why it is checkable

    Every answer is stored word for word, dated, with the sources it cited. Any figure in a report opens into the text behind it. A share you cannot trace is a claim, not a measurement.

Sampling error

What a share from five asks can honestly mean

This is the uncomfortable table, and we would rather you saw it from us than worked it out later. These are 95% confidence intervals: the range the true share could plausibly sit in, given how many times we asked.

Repeats Result Reported share Honest range
5 Named 0 times 0% 0–43%
5 Named 1 times 20% 4–62%
5 Named 3 times 60% 23–88%
5 Named 5 times 100% 57–100%
20 Named 5 times 25% 11–47%
20 Named 10 times 50% 30–70%
20 Named 15 times 75% 53–89%
20 Named 20 times 100% 84–100%

Wilson intervals, calculated rather than estimated. The practical reading: at five repeats a share tells you roughly where you stand, not precisely; at twenty it starts to be a number you can defend in a meeting.

What we do about it

We never report a share from a single ask, we say how many repeats a figure rests on, and we raise the count when a decision depends on it. Where the range is wide, the useful finding is usually not the percentage at all — it is which questions named you and which did not.

Everything that can go wrong

Seven sources of error, and the answer to each

A limitation with no answer beside it is just a worry. Each of these costs something; each has a rule that keeps it from quietly becoming your result.

Source of error How big What we do
Sampling: how many times we asked The biggest one Five repeats of a question cannot pin a share down tightly, and we say so in the table below rather than reporting 60% as if it were exact. We raise the repeat count where a decision depends on the figure, and we never report a share from a single ask.
Drift: the system answers differently anyway A few points, run to run Assistants return different answers to identical questions, with nothing changed on anybody's site. Movement under about five points is reported as no change, even when it would look like progress.
Name matching: the same business written four ways Can move a count by half Assistants write business names loosely. Merging them is a judgement, and it matters: the same run gives 63 distinct businesses counted loosely and 33 counted strictly. Unmatched names go to a review queue a person works through, and the counting rule is published with the figure.
Personalisation: the assistant knows its user Unknown, and real Chat history, account and location colour what an assistant says to a given person. We measure from a clean session with no history, which is closer to a new customer than to you — and is the reason your own ask will not match our report exactly.
Model changes under us Occasional, sometimes large Providers update models and retrieval without notice. We record which model answered each question, so a jump that coincides with a model change is visible as what it is rather than reported as your result.
Refusals: when nobody is recommended Common in health, legal, finance On some questions assistants decline to name anyone. That is recorded as a refusal, not as a zero for you — a ceiling in the category is not a failure of your business, and no supplier can sell you past it.
Missing citations Varies by assistant Some assistants show their sources and some do not. Where they do not, the source list for that answer is incomplete, and the report says so instead of implying we know what was read.

Boundaries

What we never claim

  • That we can put you in an answer

    The lists are not published, not stable and not for sale. Anyone promising placement in ChatGPT is promising something they do not control.

  • That our work caused a movement

    A report says what was changed, when, and what moved afterwards — and whether the movement was concentrated in the question group the work targeted. That is evidence, honestly labelled, not proof.

  • That one number describes you

    We do not publish a blended “AI visibility score”. Four assistants that disagree produce four numbers, and the one you are absent from is the one worth knowing about.

  • That we know what an assistant will do next

    Models and retrieval change without notice. We record which model answered and re-measure rather than predicting.

A correction we made in public

We published a number, then found it had moved

Our first pass over one measurement reported 65 businesses named across four assistants, three of them by all four. After duplicate names were merged, the same stored answers gave 63 and four. Every page now carries the new figures, and the study explains the change.

We are telling you this on the page about accuracy on purpose. A supplier whose numbers never change is not more accurate — they are just not recounting.

Read the study, with the method and the limits →

Same run, recounted

1 October 2026

First passRecount
Businesses named6563
Named by all four34
Named by one only4846

Nothing about the market changed. The counting of duplicate names did, and that is a methodological fact worth publishing.

Verification

How to check us, in four moves

  • Open any figure. Every share in the dashboard links to the answers behind it, kept verbatim and dated.
  • Ask the question yourself. Expect a different answer — then ask it five times and watch the list move.
  • Read the method. It is published in full on how it works, including how questions are written.
  • Have it repeated. Ask us to re-run a measurement, or hand the method to another agency and compare.

Common questions

Asked by people who check their suppliers

Is Sayso accurate?+

Sayso is accurate about what it measures, and the measurement has limits we publish rather than hide. What we report is a share: how often an assistant named your business out of the times we asked, per assistant and per question type, with every answer stored word for word so any figure can be traced back to the text that produced it. What we cannot do is report a precise number from a handful of asks, predict what an assistant will say tomorrow, or tell you what it says to a specific person with their own chat history.

Why did I ask the same question and get a different answer?+

Because that is how these systems work, and it is the single most common question about an AI visibility report. Assistants return different answers to identical questions, and they personalise by account, history and location. We measure from a clean session with no history, which is closer to what a new customer sees than what you see. If your own ask matched our report exactly every time, the report would be the suspicious thing.

How many times do you ask each question?+

We ask each question five times as a floor, and every figure says how many repeats it rests on. Five is enough for a share to mean something and not enough for it to be precise: a 3-out-of-5 result honestly means somewhere between 23% and 88%, which is why we raise the count when a decision rests on the figure and why we never report a share from a single ask. The table on this page gives the arithmetic in full.

What counts as real movement between measurements?+

Real movement between measurements is more than about five points on the same assistant, asked the same way. Below that we report no change, even when the direction is flattering. Reporting drift as progress is the fastest way to make the rest of a report untrustworthy.

Can I check your figures myself?+

You can check every figure yourself, and the dashboard is built for exactly that. Every share links to the answers behind it, kept verbatim and dated, with the sources each answer cited. The method is published in full on how it works, so your own team or another agency can repeat the measurement and compare.

Do you guarantee that an assistant will name our business?+

We do not guarantee that an assistant will name your business, and nobody else can either. The lists are not published, not stable and not for sale. We also do not claim that work we did caused a movement: a report says what was changed, when, and what moved afterwards, and notes whether the movement was concentrated in the question group the work targeted.

What happens when an assistant refuses to recommend anyone?+

When an assistant refuses to recommend anyone, the answer is recorded as a refusal rather than as a zero against your business. Health, legal and financial questions get hedged answers that name nobody, and that ceiling belongs to the category. Knowing which of your questions hit it is useful: it is budget you now know not to spend.

Have you ever got a published number wrong?+

Yes — we published a number that later proved wrong, and the correction is on the site. Our first pass over one measurement reported 65 businesses named, three by all four assistants. After duplicate names were merged, recomputing gave 63 and four. Every page now carries the new figures, and the research page explains the change rather than letting the number move quietly.

See the measurement on your own market

The free check runs a short version on one assistant and emails you the answers, with the raw text included. Judge the method on your own business rather than on this page.