Skip to content

How a score is produced

Everything on this page can be checked against the raw answers in the product. It is written for the person who has to defend the number, and for the person they have to defend it to.

Last updated 1 September 2026 · 12 minute read

1. What one check is

One check is one question sent to one engine once, asked as though the searcher is in the United States. The engine's full answer is stored, with the date, whether it searched the web, and every page it cited. Keylight then reads the answer for brand names. Yours, your competitors', and any others.

2. Why 25 questions and not 5

A topic like "AI writing tool" becomes about 25 different questions, phrased the way people type them. Five questions would give a score that jumps 20 points when one answer changes. Twenty-five is the point where adding more questions narrows the margin slowly enough that the money is better spent on another topic. Nothing is asked twice in the weekly run.

Figure 1 · margin of error by number of questions, at a 34% score

3. The margin of error, with a worked example

Jasper appeared in 8 of 25 answers on ChatGPT this week. The appearance score is 32%. The margin of error at 95% confidence on 25 observations is about ±18 points. The honest statement is: somewhere between 14% and 50% of the time, ChatGPT names Jasper when asked one of these questions. Keylight prints 32% ±18 and never the bare 32.

p = 8 / 25 = 0.32
margin = 1.96 × √( p (1 − p) / n ) = 1.96 × √( 0.32 × 0.68 / 25 ) ≈ 0.18
reported: 32% ±18

4. When a change counts

Engine answers vary from one check to the next. Keylight calls a change real only when the two saved results are far enough apart that ordinary variation is no longer a believable explanation. Otherwise the product says within noise. The decision comes from the answers behind those two results. Keylight does not run a separate monthly calibration or use a fixed band for any engine.

Figure 2 · two possible changes from a 28% result, with 25 answers each
Methodology | KeylightA four point rise has a range from minus twenty to plus twenty nine, which crosses no change and is within noise. A twenty four point rise has a range from plus one to plus fifty one, which stays above no change and is real.no change+4 · range −20 to +29 · within noise+24 · range +1 to +51 · real

5. The prominence reading

Appearance counts whether the brand was named. Prominence weighs each appearance by two things: how far into the answer the brand first appears, and whether the sentence around it recommends the brand, lists it neutrally, or names it as something to avoid. Each appearance scores 1.0 for a recommendation in the first third of the answer, down to 0.1 for a caution at the end. The workspace figure is the mean across appearances. It is labelled our reading everywhere, because the second part is a judgement. It is never shown on its own and never blended with appearance.

position in answer
recommends
lists neutrally
cautions against
first third
1.0
0.6
0.2
middle third
0.8
0.45
0.15
last third
0.6
0.3
0.1

6. How each engine is read

Five engines are read through the developer connection the engine's maker provides, with web search turned on. Three have no such connection for their consumer search product, so a supplier reads the engine's own web page as a signed-out US visitor would see it. The table below says which is which. Where an engine is read through its web page, the answer is what a person would have seen that day.

7. Searched, or answered from memory

Every answer records whether the engine searched the web. Some engines sometimes decide a question does not need a search and answer from what they were trained on. Those answers count toward the score, because a buyer saw them, and they are marked "the engine did not search" on the evidence page so you know the answer reflects the engine's memory, not the web this week. DeepSeek cannot search the web at all, and every DeepSeek answer carries that note.

8. The eight engines

engine
read through
searches the web
fair-use limit
Perplexity
Developer connection
Always
None
Google AI Overviews
Engine's web page
Always
None
Google AI Mode
Engine's web page
Usually
None
ChatGPT
Developer connection
Usually
None
Gemini
Developer connection
Usually
None
Claude
Developer connection
Usually
10% of capacity
Grok
Developer connection
Usually
10% of capacity
DeepSeek
Engine's web page
Never
None

Engine names and logos identify the services measured and imply no affiliation or endorsement.