Understanding how a large language model (LLM) typically talks about a category or brand is a challenge. The models are probabilistic, which means their responses vary even when a prompt stays the same.
AI marketing platforms disagree on how to solve this challenge. At Evertune, we repeat each prompt up to 100 times per model, amassing a meaningful sample size for every prompt. Other platforms argue that if you have enough unique prompts, a simple average visibility score of the entire prompt portfolio is sufficient after running each prompt only once.
For this experiment, we did both. Here's what we found.
Overall brand visibility
For our methodology faceoff, we ran 100 unique prompts about running shoes. The prompts were spread evenly across 10 topics - topics like “marathon training,” “budget picks,” “beginner runners,” etc.
When we ran prompts once each, Adidas appeared in 29% of 100 responses, with a margin of error of ±9 points. When we ran prompts 100 times each, Adidas again appeared in 29% of responses, but the margin of error shrank to ±1 point thanks to the sample size now reaching 10,000 responses.
So far, the methodologies agree: Adidas shows up about 29% of the time to prompts about running shoes. That certainly leaves room for improvement. What should Adidas do next?
Where should Adidas focus?
If Adidas wants a generative engine optimization (GEO) strategy to improve their visibility, upon which of the 10 topics should it focus its marketing time, energy and dollars?
When unique prompts are run once each, each topic has only 10 responses. Such a small sample size leaves room for response-to-response randomness to obscure meaningful signals. At 100 responses per unique prompt, each topic has a sample size of 1,000, enough data to overcome the problem of random noise.
The topic-level margin of error gets as high as ±31 points when prompts are repeated once each. It maxes out at ±3 points when prompts are repeated 100 times each.
Take “sustainable shoes” as an example. Asking each prompt once each tells Adidas that its visibility is probably between 19% and 81% (50% ±31 points). That’s not actionable. Asking each prompt 100 times tells Adidas that its visibility is probably between 51% and 57%. Now Adidas knows it has something to build on, and it has confidence that this is its 3rd-most-visible topic.
Drilling down deeper
Adidas may now want to investigate its greatest strengths and weaknesses further by drilling down from topics to prompts.
If it intends to double down on its greatest strength, “race-day shoes,” it would want to know if it is universally strong in this topic or if there are sub-topics within the topic in which its visibility differs.
Similarly, it may want to investigate if “trail running” is really such a visibility laggard or if it simply tested the wrong prompts for this topic. Because LLMs are probabilistic, slight differences in prompts can have great effects, and the longer the list of unique prompts, the more likely prompts wander away from the topic they set out to measure.
If each prompt was sampled once, then prompt visibility is binary - either Adidas was mentioned or it wasn’t. If each prompt was sampled 100 times, prompt visibility is a percentage with a margin of error of ±10 points or less.
Only repeated samples of each prompt enable Adidas to investigate these questions and drill down to the prompt level.
Repeated sampling yields actionable insights
A wide array of prompts, each sample once, can get you a clear picture of a brand’s overall visibility. It can’t do more than that.
Asking a wide array of prompts and repeatedly sampling each prompt gives you both breadth and depth. You get a clear picture of a brand’s overall visibility, and you can drill down to find topic and prompt level insights, the very insights you need to take GEO action.
Methodology
At Evertune, we track thousands of brands across LLMs by running millions of prompts a day. For this analysis, we ran 10,000 prompts on ChatGPT about running shoes, divided evenly across 10 topics. We compared the results of our full 10,000 responses to a sample of 100 responses from within these responses to compare our methodology to that of other GEO platforms.
Evertune is the AI marketing platform for brands that want to own the AI customer journey. Evertune analyzes prompt responses at scale across all major LLMs, ChatGPT, Claude, Gemini, AI Overviews and more, to deliver statistically significant visibility data, then closes the loop with tools to act on it: website optimization, data-driven content creation, most influential sources, and paid activation through affiliate and programmatic AI retargeting partners. Where most tools tell you where you stand, Evertune tells you what to do about it. Founded by early executives of The Trade Desk and backed by $20M from leading investors.

