Generative engine optimization has a measurement problem that founders keep mistaking for a strategy problem. The most complete review of the field so far examined 45 studies published between November 2023 and July 2026 and came to a narrow conclusion: content that has already been retrieved can measurably change the answer it appears in, while no technique in that literature shows a stable, cross-platform effect on whether a page gets discovered organically, or on the clicks and conversions that might follow ().
That is not a reason to skip the work. It is the reason to run it as a monitored experiment rather than a channel purchase: a set of small, conditional advantages you re-measure every month. What follows is where the evidence is strongest, how to compute your own mention share without fooling yourself, and the markup Google says you can stop worrying about.
What generative engine optimization can actually move
An AI answer is the end of a pipeline, and every stage can fail on its own: the assistant decides whether to search at all, retrieves a handful of documents, ranks them into its context window, uses one or two, and maybe attaches a link. Most of what gets sold as GEO acts on one stage and is silent about the rest.
The paper that named the field, Aggarwal and colleagues' "GEO: Generative Engine Optimization" (KDD 2024), worked on the middle of that pipeline. It took 10,000 queries, gave the top five Google results for each to GPT-3.5-turbo, then rewrote one of those five documents using strategies such as adding quotations, statistics or citations. The best-performing strategies raised that document's position-adjusted share of the answer by 30% to 40% ().
Read that result for what it is. Every document was already inside the model's context. The experiment measured what happens after retrieval, which is why the survey notes that the widely repeated "up to 40%" says nothing about whether more readers will click, or whether your page becomes more likely to be retrieved at all.
The click stage is documented more soberly. In Pew Research's study of 68,879 Google searches from March 2025, users who saw an AI summary clicked a traditional result in 8% of visits, against 15% when no summary appeared, and only 1% of visits involved clicking a link inside the summary itself (). Google's own guidance argues the opposite half of the same picture: clicks arriving from pages with AI Overviews are "higher quality", meaning people tend to spend more time on the site (). Fewer clicks, more deliberate ones, is the workable reading.
"GEO SEO" and "AI search optimization" are one job under three names
The three phrases show up in the same search results and describe the same work. The difference is who says them. "Generative engine optimization" comes from the research literature. "GEO SEO" is the compressed version founders type, and it collides with geographic and local targeting, so it is worth confirming which one a vendor is actually selling you. "AI search optimization" carries the buying intent, which is the phrase people use when they are comparing tools and budgets.
For scale, US search demand in September 2026 ran about 4,400 searches a month for "generative engine optimization", 2,900 for "geo seo" and 1,300 for "ai search optimization", the last of the three at the lowest difficulty of the group (DataForSEO, pulled for this site). Those are three ways of asking one question, not three markets to attack in sequence.
The practical split is in what each reader wants. Informational searchers want the mechanics; buyer-intent searchers want the checklist and the price, and any article aimed at them has to survive one test: what does the tool actually measure? A vendor reporting sampled mentions is describing visibility. A vendor reporting revenue from AI answers is claiming attribution that no study has established yet.
Start with the questions buyers type, not the keywords you want to rank for
Long, question-shaped queries are where AI answers actually appear. In Pew's dataset, 8% of one- and two-word searches produced an AI summary, against 53% of searches with ten words or more and 60% of queries beginning with who, what, when or why. A later audit of 55,393 trending queries put the overall activation rate at 13.7%, rising to 64.7% when the query was phrased as a question (Xu et al., summarized in the ).
This is the pragmatic core of the work for a small product. The founder complaint that AI assistants "recommend the same 5-10 big names in every category", as one r/microsaas commenter put it in April 2026, holds at the category level. A preprint testing 112 startups found ChatGPT named a product correctly 99.4% of the time when the product was named, but surfaced it in only 3.32% of organic discovery queries, with Perplexity falling from 94.3% to 8.29%. Being known and being chosen are different states, and the second one is decided by the question.
So build a question list instead of a keyword list, and pull it from places where the questions already exist: your support inbox, your last ten sales conversations, and the threads where buyers ask about your category. Twenty to forty questions is plenty. Sort them by what the asker wants: what is X, X versus Y, which X fits [situation], how do I fix [problem].
Then decide where each answer belongs. Definitional and comparison questions can live on your own site. Community questions generally belong on the community's surface, where you disclose that you build the thing you are recommending. Those surfaces also appear to carry weight with the assistants: Pew found Wikipedia, YouTube and Reddit were the most frequently cited sources in AI summaries, 15% of listed sources collectively. You do not control them, which is exactly why you measure instead of assuming. RiseMore's prompt discovery exists for this step and its discovery is English-only, which matters if your buyers ask in another language.
How to tell whether you are in the answer
A single sample is noise. Schulte et al. measured four engines across 45 days and found daily source-level overlap (Jaccard) of roughly 0.34 to 0.42, with repeats inside 24 hours scoring about the same, and suggested seven to eight repetitions per prompt as a starting point. Kirsten et al. found page overlap across two months of 18% for AI Overviews against 45% for organic Google, and repeated runs at temperature zero changed 9% to 28% of decisions. Between surfaces, Grossman et al. reported URL-level Jaccard of 0.11 to 0.18 across organic Google, AI Overviews and Gemini (all summarized in the ).
A protocol that survives that noise, runnable in a few hours a month:
- Pick 15 to 25 buyer questions, the specific ones, not your category name.
- Write three paraphrases of each, because small reformulations change which sources get used.
- Run each paraphrase five to eight times per assistant, on the same day, logged in and logged out if you have both.
- Record every outcome, including the runs where the assistant did not search. In Schulte's setup, 57.8% of ChatGPT repetitions never activated web search; a run with no citations is not a miss for you, it is an answer that never consulted the web.
- Report two numbers, not one: how often you are named, and how often you are cited with a link. Those are different outcomes with different causes.
- Compare month to month. Daily movement in a sampled answer is mostly the sampler.
The arithmetic looks like this. Twenty questions, three paraphrases, five runs each gives 300 answers. A product named in 24 of them sits at an 8% mention share. If 11 of those 24 mentions came from runs with no web search, the number you can act on is smaller, and blending the two is the most common way founders overstate their visibility.
RiseMore's GEO agent runs one discovery and monitoring pass a week (Wednesdays, 02:00 UTC) across ChatGPT, Gemini and Perplexity, and reports AI visibility as the share of sampled answers that mention the product, with 30- and 90-day trends. Two limits matter before you trust the chart. Monitoring stops when credits run out, so a gap in the line is a billing event rather than lost visibility, and a mention is a visibility sample rather than a recommendation. The agent is part of the paid tiers (Pro, $20/month, metered in credits rather than post counts); the free plan does not include it.
The levers with the most support, and the rewrite that backfires
Relevance and position come first. Wan et al.'s counterfactual experiments found models favour explicit alignment with the question, often over human credibility cues such as scientific references or a neutral tone. Puerto et al. found that moving a source higher in the context had a larger effect than most rewrites. Vishwakarma et al.'s factorial experiment, 252,000 trials across six models and eighteen factors, identified relevance and position as the primary determinants of who gets the first citation. In practice: answer the question in the opening sentences of the section that addresses it, and accept that a page which never gets retrieved cannot be cited however well it is written.
Extractable evidence comes second. Prices, dates, definitions, comparisons and references are units a model can lift and attribute; Vishwakarma et al. found effects for explicit prices and recent dates, while formatting changes on their own did little. The test for the writer is whether the numbers are ones you can stand behind, because the same signals raise reuse and degrade the answer when a statistic is invented, and the survey's integrity check asks whether your references are verifiable.
Then the counter-evidence, which is the part the marketing posts leave out. Puerto et al.'s C-SEO Bench tested 54 method and domain combinations across roughly 1,900 queries and 16,360 documents: three were significantly positive in the main experiment, none was positive for question answering, and gains shrank as more sites adopted the technique. When Kim et al. put retrieval and reranking back into the pipeline (SAGEO Arena, 171,003 documents, 2,700 queries), body-only optimization cut average presence in the top 20 by about 9%, top-10 presence after reranking by 16%, and final citation by 6%. A page rewritten to be quoted can become harder to retrieve.
Where the gains do appear, they are redistributive and kinder to the outsider. In the original paper's Cite Sources condition, the fifth-ranked source gained 115.1% while the top-ranked source lost 30.3%. Break into a category where the assistant already names five established tools and the marginal document has more to gain than the incumbent has to defend. The same finding says your advantage decays as competitors do the same work, which is an argument for measuring rather than endlessly rewriting.
Schema, FAQ markup and llms.txt: what Google says you can skip
Google's documentation is direct on this: to appear in AI features you do not need new machine-readable files, AI text files or markup, and "there's also no special schema.org structured data that you need to add" (). Its optimization guide goes further and lists what it says you can ignore for Google Search: llms.txt and similar files, "chunking" content into small pieces, rewriting text specifically for AI systems, and pursuing inauthentic mentions ().
Structured data still earns rich results, and product and business details remain worth keeping accurate. That makes markup a poor place to start when time is the constraint, because no tag substitutes for being one of the documents the model reads.
Keep the boundary of that advice in view. These are Google's rules for Google. ChatGPT, Gemini and Perplexity publish no equivalent, and the absence of a shared cross-engine lever is one reason the survey found no technique with a stable effect across platforms.
What to do with the next month
Two or three genuine answers a month is a workable cadence. Take them from your question list, write the version that settles the question better than the current top five results do, put your prices, dates and the comparison you can defend into it, and publish on the surface where the question is asked. Then run the sampling once a month and write down one number.
The budget is less mysterious than it looks. The writing costs time, the measuring costs either the same time or a paid tier (RiseMore bundles prompt research, drafting and weekly GEO monitoring into Pro at $20/month, where credits meter everything the agents do, so a plan is not a fixed number of posts). What no tier buys is a promise.
The traffic claim is where the evidence is thinnest. One log-based study of a site where some pages received an AEO intervention saw total ChatGPT referrals rise by a factor of 5.7, while untreated pages rose by 3.5 over the same period as the platform grew; the controlled estimate of the extra lift was 1.82 with a wide interval, and a placebo test did not reach significance. Treat any revenue forecast built on AI answer visibility as a hypothesis, and check it against the one number you can actually observe about your own product.
Before you publish anything, write down your current mention share. In ninety days that number is the only honest answer to whether the work moved, and the question list you build to get it keeps paying for itself in sales calls and site copy even if the answer is no.



