Contents
Why the same question gets a different answer
Other people have measured this too
What an unqualified score can't tell you
What a margin of error actually means
How Keylight handles it
Define
·
6 minute read
·
Why do AI answers change every time you ask?
These systems sample an answer at the moment you ask, rather than looking one up. Ask the same question twice and you can get two different answers, so a visibility score taken from a single run is one draw from a distribution, not a settled measurement.
That single fact is why a visibility score needs a margin of error printed beside it, and why a score quoted without one is telling you less than it looks like it is telling you.
Why the same question gets a different answer
Four things are moving underneath a single question, and a visibility score reports none of them.
The engine is not retrieving a stored fact when it answers. It builds the answer one word at a time, and at each step it picks from a weighted range of plausible next words. There is no single correct one sitting in a database to fetch. Two runs of the identical prompt can diverge on the first sentence and compound from there.
Retrieval moves independently of that. Perplexity, Google AI Overviews and, when it decides to search, ChatGPT each run a fresh web search behind the answer, and the pages that search returns are not fixed. A page ranking third an hour ago can rank fifth now, or drop off the results the engine sees at all, which changes which brands the engine has in front of it to quote from.
Where the question comes from changes what comes back too. Google's Gemini has no setting for the location of the person searching, so it answers from wherever the request happens to come from. Sending it the identical marketing question from Delhi returned Indian sources, recorded in Keylight's own testing on 7 September 2026. A US agency and a freelancer in another country typing the same question into Gemini are not necessarily shown the same web.
And the engine itself is not the same engine from one week to the next. Providers update the model answering a question without announcing it, so a score that moved because the model changed underneath it is not the same as a score that moved because the brand did.
Other people have measured this too
An academic survey of the field is the clearest outside confirmation, because it did not set out to make a point about visibility tracking and arrived at one anyway. Olivier Martinez's critical survey of 45 studies on how brands try to shape what generative AI engines say about them, submitted to arXiv on 15 July 2026, looked specifically at published audits of commercial AI engines and found "low source overlap, substantial run-to-run variability, and persistent fidelity gaps." Across the full set of 45 studies, not just those audits, the survey's own synthesis is blunter: "no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior."
Search Engine Land reported the study on 28 January 2026. It was run by Rand Fishkin of SparkToro and Patrick O'Donnell of Gumshoe.ai, which sells AI visibility tracking: "Six hundred volunteers ran 12 identical prompts through ChatGPT, Claude, and Google's AI nearly 3,000 times." Getting the identical brand list back twice happened fewer than once in a hundred attempts. Getting it back in the identical order happened closer to once in a thousand.
A separate test at Washington State University asked ChatGPT the same question ten times in a row, worded identically each time. Lead researcher Mesut Cicek described the result in the university's own account of the study, published 16 March 2026: "We used 10 prompts with the same exact question. Everything was identical. It would answer true. Next, it says it's false." The same account reports that across those ten identical prompts ChatGPT judged the statements accurately 73% of the time. That figure measures how often it was right, not how often it repeated itself. The flip from true to false, on a question that had not changed, is the part that matters here.
What an unqualified score can't tell you
None of this makes a visibility score useless. It makes a score with no margin of error impossible to act on, because there is no way to tell a real change from the engine's own wandering.
A business that checks its visibility today and gets 30%, then checks again two weeks later and gets 36%, has not learned that anything improved. Keylight builds a weekly score from 25 questions per topic, and at 25 questions a score near 30% carries a margin of about 17 points either way, so a six point move sits well inside it with nothing about the brand having changed at all. Reported without a margin of error, it reads as progress worth reporting to a client. It might be six points of static.
What a margin of error actually means
A margin of error states how far a single score could plausibly sit from the truth, given how many questions it is built from. More questions narrow the range. Fewer questions widen it, and the widening is not gentle.
Keylight got this number wrong before it got it right. The formula usually taught first, based on a normal approximation, breaks exactly where a visibility score lives. On a brand named in every one of 20 answers, that formula reports a margin of error of zero, which cannot be true. On a brand named in 19 of 20, it puts the upper bound above 100%, which is impossible. Both happen routinely at the sample sizes this category runs at. Keylight used that formula in its own code until 5 September 2026, then switched to the Wilson interval, a different way of calculating the same range that does not fail in either direction.
The margin depends on the score as well as the sample, and it is widest at a score near half. Take that widest case and watch what sample size does: about plus or minus 33 points on 5 questions, about 18 points on 25, and about 10 points on 100. The 17 points quoted earlier is that same 25 questions at a score near 30%, where the range is a little narrower.
A score built from five questions is barely a score. A score built from a hundred is a real number worth reading. Keylight sits at 25 per topic. That leaves a margin wide enough that small weekly moves mean nothing, which is exactly why it refuses to report them as changes.
How Keylight handles it
Every score Keylight reports carries its margin of error beside it. A week-on-week move is only called a change when its whole range sits clear of zero, so a difference the sample cannot tell apart from nothing is not reported as one. The answer each engine gave is kept behind that engine's score, on an evidence page, so anyone reading a number can go and read the words it came from.
What the industry calls moving these numbers is generative engine optimization. Its most-proposed technique is tested in what the evidence says about llms.txt, and the tools selling this measurement are compared in the best AI visibility tools.
Glossary
Appearance score
Prominence reading
Noise band
Discovered brand
Cited page