Contents
Where the term came from
What a generative engine does when it answers
What the evidence supports and what it does not
How you would know whether it worked for you
Define
·
6 minute read
·
What is generative engine optimization?
Generative engine optimization, usually shortened to GEO, is the attempt to change how often and how favourably a brand or a page shows up in the answers written by AI systems such as ChatGPT, Google's AI Overviews, Perplexity, Gemini and Claude. The evidence behind it is thinner than the number of people now selling it.
Where the term came from
GEO was named in a paper called "GEO: Generative Engine Optimization", posted to arXiv in November 2023 by Pranjal Aggarwal and five co-authors, and accepted to the ACM SIGKDD conference in 2024. The paper built a benchmark called GEO-bench and tested rewriting web content, adding things like statistics, quotations and citations, to see whether an AI-written answer used that content more. Its headline finding, on the paper's own arXiv listing read 20 September 2026, is visibility improvements of up to 40 percent. The detail about which rewrites did the work is in the full text rather than that listing, which says only that the benchmark exists.
That 40 percent figure is the headline number from the 2024 paper. The condition attached to it is the most important thing on this page. A July 2026 academic survey of 45 studies on generative engine optimization found that the 2024 paper's gains are, in its own words, "valid within its experimental setting but conditional on a source already being present in a fixed context" and that they "establish neither organic discoverability nor durable traffic effects." In plain terms, the original study measured what happened once a page had already been picked as one of a small number of sources an AI was going to read from. It did not test whether any rewrite made a page more likely to be picked in the first place, and it did not measure whether any of it sent a single visitor anywhere. Getting chosen and getting described well once chosen are two different problems, and the widely quoted number answers only the second one.
What a generative engine does when it answers
A generative engine does not hold a ranked list of web pages the way a search engine does. When it needs current information to answer a question, it runs a search of its own, pulls back a small set of pages or snippets, and hands them to a language model as background reading. The model then writes an answer in its own words, choosing which of those sources to name, quote or recommend, and in what order. A brand cannot be mentioned in that answer unless one of its pages made it into that small set of background reading in the first place, and being included is no guarantee the model chooses to use it.
Not every answer works this way. Some questions get answered from what the model already learned during training, with no search and nothing to cite. Whether a given answer involved a search at all changes what, if anything, a brand could have done to appear in it.
What the evidence supports and what it does not
The most useful check on any of this is not one company's claim but an academic review. Olivier Martinez surveyed 45 studies on generative engine optimization published between November 2023 and July 2026, posted to arXiv on 15 July 2026 as "Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)". Read from the paper's abstract on 16 September 2026, four findings hold up across the studies reviewed.
Topical relevance and where a source sits among the material the model reads are "the most reproducible levers": content that matches the question closely and appears early in what the engine reads tends to get used more consistently than content that does not. "[G]eneric heuristics transfer poorly": a rewriting trick that worked in one study, on one engine, for one kind of question, is not reliable evidence that it works anywhere else. "[C]itation-oriented rewrites can impair retrieval": loading a page with citations and statistics specifically to look more quotable can make an engine pull it up less often, not more. And "competition can erode individual gains": once everyone writing in a category adopts the same trick, the edge any one of them gets from it shrinks or disappears.
Across all 45 studies, the survey's conclusion is direct: "no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior." Nothing reviewed has been shown to reliably get a brand chosen more often, over time, on more than one engine, in a way that changed what happened next.
None of this means nothing matters. It means the evidence, read as of the most recent academic review of it, covers a narrow claim: a source an engine is already reading can sometimes be made more likely to be quoted from. Whether a source gets read by the engine in the first place, whether that reading turns into a visit or a sale, and whether any of it still holds next month or on a different engine, is not something the reviewed research has shown.
How you would know whether it worked for you
Since nothing reviewed shows a reliable technique, the practical question moves from which tactic works to how anyone would even tell. The same survey reports that published audits of commercial AI engines, meaning academic work examining the engines rather than products sold to track brand visibility, show "substantial run-to-run variability": ask an engine the same question on different days, with nothing about the brand having changed, and the answer moves on its own. Read from the survey's abstract on 16 September 2026.
That fact alone rules out judging a change by a week's difference in results. A brand's mentions going up on Tuesday and down on Thursday could just as easily be the engine's ordinary variation as evidence that anything worked or failed. Telling the two apart needs enough separate questions behind each reading that the range around it is narrow, and enough readings over enough weeks that a move can be told from the ordinary wobble.
Keylight, the tool this page sits on, is built for that kind of measurement. It asks a fixed set of questions on a schedule, tracks how often a brand's name appears in the answers, and reports a margin of error next to the number, calling a week's movement a change only when its whole range sits clear of zero. That is measurement, not a claim that any tactic caused what it recorded. Nothing above shows that it will, and nothing here should be read as saying otherwise.
Three separate published measurements of how far AI answers move between identical runs are set out in why AI answers change between runs. The one technique most often proposed in this space, publishing an llms.txt file, is held against the evidence in what the evidence says about llms.txt, and the tools sold to measure any of this are compared in the best AI visibility tools.
Glossary
Appearance score
Prominence reading
Noise band
Discovered brand
Cited page