TL;DR
- Evertune samples every prompt up to 100 times per model. Other GEO platforms only sample each prompt once per model.
- Sampling prompts once each fails to uncover 9 out of 10 of the source URLs models use to form their responses, compared to sampling each prompt 100 times.
- Source insights at the topic and prompt levels are lost when prompts are sampled once each. The importance of individual domains becomes highly random and variable.
Large language model (LLM) responses depend on many things, especially foundational knowledge learned in training, supplemental information retrieved via search, the cadre of publishers and brands that give bots permission to access their content and the model-specific behaviors of each LLM and LLM version.
There is also an element of randomness, since LLMs are probabilistic. Responses change even when the prompt stays the same, making it essential to repeatedly sample prompts for statistically meaningful insights.
There are two primary ways generative engine optimization (GEO) platforms address this statistical noise. At Evertune, we repeat each prompt up to 100 times per model, amassing a meaningful sample size for every prompt. Other platforms argue that if you run enough unique prompts, sampling each prompt once is sufficient.
We have previously shown how this methodology can misinform brands at the topic level. Inherently large margins of error obscure a brand’s true visibility at the topic level and make it impossible to monitor the success or failure of GEO efforts.
We now turn our attention to another major shortcoming of sampling prompts once each: the inability to reliably identify the sources LLMs use to form their responses. This is exactly what brands need to know before adding and refreshing content to get cited by AI.
Small samples miss most source opportunities
We ran 100 unique prompts about running shoes across ChatGPT, Gemini, Google AI Mode and Google AI Overview, and for every model, sampling each prompt once missed about 90% of the cited URLs that were included when each prompt was sampled 100 times.
LLMs search far and wide when they incorporate search results into their responses. When prompts were sampled 100 times each, each prompt sourced from about 60-210 unique URLs depending on the model, with ChatGPT the low and Google AI Mode the high.
But within that wide array of URLs, certain URLs were cited more than others. Some of these consistently cited URLs were missed when each prompt was sampled only once - including ChatGPT’s 8th-most-cited link and Gemini’s 7th-most-cited link.
Even at the domain level, sampling each prompt once leaves each site’s source share with a large margin of error that makes it difficult for a brand to understand which content hub to target. For example, the source share for runnersworld.com - ChatGPT’s most-cited domain for this set of prompts - falls somewhere in the range of 15-20% based on the 1-response-per-prompt method. With a margin of error of 2.5 points, you can’t get more precise than that 5-point spread. At 100 responses per prompt, the margin of error is only 0.3 points, yielding a precise source share of 17.1-17.7%.
These key shortcomings - not capturing highly cited links and lacking precision about sources’ importance - are significantly compounded when you try to dig below the overall surface to get insights at the topic and prompt levels.
Source insights lost to randomness
We previously demonstrated that the practices of repeatedly sampling each prompt and running each prompt once can agree on overall figures, but sampling each prompt once fails when you want to drill down into topics and prompts. The same is true when it comes to understanding which sources shape AI models’ responses.
In this experiment, we ran 100 running shoe prompts equally distributed across 10 topics. The frequency with which any domain was cited varied significantly by topic. For example, ChatGPT cited Runnersworld.com 30% of the time in the “beginner runners” topic but only 5% of the time in “sustainable shoes.”
That kind of topic level insight is exactly what a brand needs to know to get the most bang for its buck in a GEO campaign, and they are the insights that are obscured by statistical noise when topic-level sample sizes are small.
Say that a brand wanted to concentrate on the topic “marathon training.” If their goal is to get cited by AI, which domains should they prioritize? We ran 100 iterations of one response per prompt to demonstrate how much the answer to that question can vary depending on that particular set of responses.
Across our 100 iterations, seven different domains appeared as ChatGPT’s most-cited for the “marathon training” topic at least once. Irunfar.com, which is the third-most cited domain at 100 repetitions per prompt, went completely uncited in two iterations. Reddit’s rank varied wildly, spanning from 1st to 23rd.
The where and how of GEO
Sampling each prompt repeatedly is the only way to reliably identify both where your brand’s visibility is weak and where to target content to improve it. A meaningful sample of responses tells you which topic to focus on, and the sources most cited in those responses tell you which opportunities - affiliate, owned, earned, etc. - exist for you to become more visible.
Methodology
At Evertune, we track thousands of brands across LLMs by running millions of prompts a day. For this analysis, we ran 40,000 prompts on ChatGPT, Gemini, Google AI Mode and Google AI Overview about running shoes, divided evenly across 10 topics. We compared the results of our full 10,000 responses per model to repeated 1-response-per-prompt samples, which simulate the methodology of other GEO platforms.
Evertune is the first Generative Engine Optimization (GEO) platform built to explore, measure, act and advertise across the entire AI customer journey, connecting brands directly to ChatGPT Ads and programmatic advertising partners like The Trade Desk and Index Exchange. Where most GEO tools sample each prompt once a day, Evertune samples every prompt 100 times per model across 11+ AI models, delivering statistically significant visibility data at half the cost of competing platforms. Evertune's agents do the work for you: a prompt agent mines over 150 million real user conversations to find the exact questions buyers ask, an insights agent tells you what to do next each week, and an ads agent builds a complete ChatGPT campaign around your visibility gaps. From there, Evertune closes the loop with website optimization, data-driven content creation, source-level influence mapping and paid activation through affiliate and programmatic AI retargeting partners. Founded by the team that pioneered programmatic advertising at The Trade Desk, now building the next marketing channel.

