Most AI search competitive analyses are one person typing five prompts into ChatGPT on a Tuesday afternoon and reporting what came back as though it were a market finding. It is not. A single run of a handful of prompts on one engine tells you what happened once, not what is typical. A benchmark that actually holds up needs a fixed method, repeated over time, with enough prompts to smooth out noise.
Why a screenshot is not a benchmark
AI engines do not return the same answer to the same prompt every time. Model updates, retrieval changes, and even ordinary variance in generation mean that running one prompt once and screenshotting the result captures a moment, not a pattern. Treating that moment as evidence of who is winning in AI search is the same mistake as calling a single day's stock price the company's valuation. The fix is not a better prompt. It is more prompts, run more than once, logged consistently.
Building the prompt set
Group prompts by buying stage rather than writing them as a flat list. Problem-aware prompts describe a pain without naming a category, 'why does our sales team miss follow-up on inbound leads'. Comparison prompts name a category and ask for options, 'best tools for routing inbound leads to sales'. Vendor-specific prompts name your company or a named competitor directly, 'is [competitor] good for a 50-person sales team'. Each stage tends to surface a different set of winning content types, so keep them separate in the log rather than blending them into one score.
Pull the actual language from sales call notes and support tickets where possible, since that is closer to how buyers phrase problems than anything a marketing team would draft from a keyword tool. A prompt set built entirely at a desk tends to sound like search engine optimisation, not like a person asking a question.
Choosing the engines and the sampling method
Resist the urge to cover every engine on the market. Two or three, chosen based on where your buyers actually are, produces a benchmark you can maintain consistently. A wide net across eight engines sampled once beats nothing, but a narrow net across three engines sampled every week for a quarter produces a far more defensible trend, and it is the trend that a leadership team should actually care about, not any single data point.
For each run, use a logged-out or fresh session where the engine allows it, keep the query wording identical to the last run, and note the date and any known model updates that occurred since the previous run. That last detail matters more than it sounds: a citation share swing that lines up with a publicised model update is a different finding than one that lines up with nothing, and conflating the two leads a team to credit its own content work for a shift caused entirely by the platform.
Worked example
Illustrative modelA one-month competitive benchmark
Illustrative model. A B2B project management software vendor runs a 50-prompt set across ChatGPT and Perplexity, once a week for four weeks, against two named competitors. Figures are assumptions to illustrate how the analysis should be read, not a reported outcome.
- Problem-aware prompts (18 of 50)
- Own domain cited on 2, Competitor A on 9, Competitor B on 4
- Comparison prompts (20 of 50)
- Own domain cited on 6, Competitor A on 11, Competitor B on 7
- Vendor-specific prompts (12 of 50)
- Own domain cited on 10, Competitor A on 8, Competitor B on 3
- Content type behind Competitor A's wins
- 14 of 20 citations traced to a single comparison page updated monthly
Result: The aggregate citation count makes it look like the company is losing broadly. The segmented view shows it is actually winning vendor-specific prompts and losing almost entirely on one comparison page that a competitor keeps current, which points at a single, specific page to build rather than a vague mandate to 'do more AI search'.
Reading the gap without overreacting to it
The strongest case against this
If a competitor dominates citation share across most prompts, the honest conclusion is that they have a structural advantage, more domain authority, more content volume, more years of indexing, and no amount of page-level fixes will close that gap.
Domain-level advantages are real and worth naming rather than ignoring. But citation behaviour in generative engines rewards specific, current, well-structured answers more than it rewards sheer site size, which is different from how traditional organic ranking works. A segmented benchmark usually finds pockets, often comparison or vendor-specific prompts, where a smaller company with one well-built page can win a citation a much bigger competitor is not defending closely. The right response to a big structural gap is not to give up on the whole category, it is to stop competing on the broad prompts where the gap is real and concentrate on the narrower prompts where a specific page can actually move the number.
Turning the benchmark into a brief
A benchmark that ends as a slide with a citation share percentage on it has not done its job. The output should be a short, ranked list: which prompts are being lost, to which competitor, on which content type, and what a page would need to include to compete for that citation. If the log shows a competitor consistently wins because their page states pricing directly and yours requires a form fill to see it, that is a specific, buildable fix, not a strategic mystery.
Common mistakes when building the prompt set
The most frequent error is writing prompts that describe the product category rather than the buyer's actual problem. A prompt like 'project management software' is close to a keyword, not a question, and it tends to produce generic, high-authority results that tell you little about where a smaller player could realistically compete. A prompt like 'how do I get visibility across three teams without adding another dashboard nobody checks' is closer to how a real buyer thinks, and it produces citation results that are far more useful to act on because the page that wins it had to actually address the underlying situation.
A second mistake is letting one person write the entire prompt set from their own assumptions about the buyer. Prompt lists built this way tend to cluster around whatever that person finds interesting or has heard recently, missing whole segments of the buying journey. Pull prompts from multiple sources, sales call transcripts, support tickets, win-loss interviews, and customer success notes, so the set reflects the actual range of questions being asked rather than one person's mental model of the market.
- Writing prompts as short category keywords instead of full buyer questions.
- Building the entire list from one person's assumptions rather than sales and support source material.
- Skipping vendor-specific prompts that name real competitors, which are often the easiest citations to win or lose visibly.
- Letting the prompt count balloon with near-duplicate phrasing that adds noise instead of new signal.
- Never revisiting whether the prompts still match how the market talks a year later.
Presenting the benchmark to a leadership team
A benchmark that only lives in a spreadsheet does not change a budget conversation. When presenting this upward, lead with the segmented finding rather than the aggregate number, since an aggregate citation share on its own invites a simple, often demoralising read that a segmented view usually complicates in a useful way. Showing that the company loses broadly on problem-aware prompts but wins narrowly on vendor-specific and comparison prompts gives a leadership team something to act on: where to defend, where to concede for now, and where a specific page investment has a realistic chance of moving the number.
Resist the pressure to turn one month of data into a definitive verdict, especially in a room that wants a clean answer. State plainly that the benchmark is a repeatable process, not a one-time score, and that the next month's run is what will show whether a specific investment is paying off. A leadership team that understands this going in is far less likely to overreact to a single noisy data point later, and far more likely to keep funding the tracking work that makes the whole exercise credible.
Before presenting a competitive benchmark upward
- ✓The finding is segmented by buying stage, not a single aggregate percentage.
- ✓Every cited competitor page is linked to a specific, buildable fix, not a general observation.
- ✓The deck states clearly how many weeks of data back the finding, and does not present one run as a trend.
- ✓There is a named next action and owner for the highest-priority gap, not just a report.
- ✓A re-test date is already on the calendar before the meeting happens.
Automating parts of the benchmarking workflow
Running 40 to 60 prompts across two or three engines by hand every week is real effort, and most teams eventually want to automate at least the collection step. Several third-party visibility platforms exist specifically to run a fixed prompt set on a schedule and log citations automatically, which removes the manual burden once the prompt list is stable. The tradeoff is the same one that applies to any vendor tool in this space: you are trusting their sampling method, their definition of a citation, and their update cadence, so ask to see the raw prompt-level data before treating any dashboard summary as the finding itself.
A middle path that works for many teams is to automate the collection but keep the segmentation and interpretation manual. Let a tool or script handle running the prompts and recording raw citations, then have a person tag each citation by content type and buying stage and write the resulting brief. This keeps the labour-intensive part automated while keeping the judgment calls, which content type won and why, in human hands where they belong for now.
Handling competitors who game the benchmark, or seem to
Occasionally a benchmark will surface a competitor whose citation share looks implausibly high across almost every prompt in a category, and it is worth pausing to check whether that reflects genuine content strength or an artefact of the sampling method before reacting to it. Check whether the citations are landing on genuinely different pages addressing different angles of the topic, which suggests real content depth, or whether the same one or two pages are being cited repeatedly across many different prompts, which suggests the engine has simply decided that domain is a default trusted source for the category rather than that each individual page is winning on its own merits.
This distinction matters for how you respond. A competitor winning through genuine, deep, page-by-page content coverage is a harder problem to compete with directly, and the honest response is often the segmented approach described earlier, conceding the broad prompts and concentrating on narrower ones. A competitor winning through a kind of default domain trust rather than depth is a softer target, because a single, sharply better page on a high-value prompt has a real chance of displacing a citation that was never that strongly earned in the first place. The benchmark log, read carefully at the page level rather than just the domain level, is usually enough to tell the two situations apart.
How competitive benchmarking should change your content roadmap
A competitive benchmark that never changes the order in which pages get built is a benchmark that exists for its own sake. The most direct way to connect the two is to treat the ranked list of lost prompts, tagged by which competitor and which content type won them, as a standing input to the content roadmap review, sitting alongside keyword research and sales requests rather than as a separate, occasional exercise that lives in its own deck. When a prompt has been consistently lost to the same competitor page for two or three benchmark cycles in a row, that page should outrank almost anything else competing for the next content sprint, because the cost of continuing to lose it compounds every month it stays unaddressed.
It also helps to separate genuinely new content from targeted rewrites when planning off the benchmark, since the two require different effort and produce different kinds of results. A prompt that is lost because no page on your site addresses the topic at all calls for new content, and the benchmark's segmentation by buying stage tells you what kind of page to build first, a comparison page for a comparison-stage loss, a documentation-style page for a technical one. A prompt that is lost despite an existing page addressing the topic calls for a rewrite, and the specific gap, missing pricing, vague claims, weak structure, should be named in the brief rather than left for the writer to guess at.
Finally, resist the temptation to chase every lost prompt at once. A benchmark with forty or fifty prompts will usually surface more gaps than a content team can address in a single quarter, and treating the whole list as equally urgent guarantees that nothing gets finished properly. Rank the gaps by a combination of how often the prompt type comes up in real buyer conversations and how close the existing page already is to competitive, and commit to closing a small number of them completely before moving to the next batch, rather than making shallow progress across all of them at once.
What to do with the trend over time
The value of running this monthly rather than once is that it turns a snapshot into a trend, and a trend is what tells a leadership team whether the work is moving anything. A single month's citation share number invites the wrong question, are we winning or losing. A four-month trend, tied to specific page changes shipped in between, invites the right one: did this specific change move this specific number, and is it worth doing again on the next page in the queue.
Set up the benchmark
- 1Write 40 to 60 prompts grouped by buying stage: problem-aware, comparison, and vendor-specific.
- 2Pick two or three engines to sample, based on where your own analytics show AI referral traffic actually coming from.
- 3Run the full prompt set on a fixed day each week and log domain, content type, and link presence for every citation.
- 4Tag every competitor citation by which content type won it, so the gap analysis points at a specific page to build.
- 5Re-run the identical prompt set for at least four weeks before drawing a conclusion from the trend.
Common questions
- How many prompts do I need for a competitive benchmark to be meaningful?
- Forty to sixty prompts is a workable range for a single product category. Fewer than that and a single engine update can swing the whole result. More than that becomes expensive to run and log by hand without automation, and most teams do not have distinct enough buying questions to fill a much larger set without padding it with near-duplicates that add noise rather than signal.
- Which AI engines should be included in the benchmark?
- Start with whichever engines already show up in your own analytics as referral sources, since that tells you where your actual buyers are asking questions. If you have no referral data yet, ChatGPT, Perplexity, and Google AI Overviews together cover the large majority of B2B AI search behaviour and are a reasonable default starting set.
- Should the prompt set match my existing SEO keyword list?
- No, and this is the most common mistake in this exercise. SEO keyword lists are built around short, high-volume search terms. AI engines respond better to full questions and comparisons, closer to how someone would ask a colleague. Write the prompt set fresh, based on the actual questions your sales team hears on calls, rather than repurposing a keyword spreadsheet.
- How do I compare citation share fairly against a much bigger competitor?
- Segment the prompt set by buying stage and compare within each segment rather than only looking at an aggregate score. A large competitor often dominates broad, problem-aware prompts through sheer content volume and domain authority, but a smaller company can win a disproportionate share of narrow, comparison-stage prompts where a more specific and current page beats a generic one. The aggregate number can hide that split.
- What should I record for each citation, not just whether one occurred?
- Record the domain cited, the specific URL if visible, the content type of that page, comparison page, documentation, review site, or blog post, and whether the citation carried a clickable link or was a text mention only. This level of detail is what turns a benchmark into a brief. Knowing that a competitor's comparison page wins a prompt is actionable. Knowing only that they were cited somewhere is not.
- How often should the competitive benchmark be re-run?
- Monthly is a sensible cadence for most B2B teams, run against the identical prompt set so the trend is comparable. Running it more often adds tracking overhead without adding much insight, since AI engine behaviour and your own content changes both take time to show a stable pattern. Running it less often risks missing a competitor's content push before it has already won several quarters of citation share.
Last updated
- First published
- Last updated
- Last fact review
Who wrote this
Avishai Sam Bitton
Founder, DemandBox
Avishai runs demand generation programs for B2B SaaS companies across performance marketing, SEO, and answer engine optimization. He works directly with the teams he advises, with no account managers in between.
Connect on LinkedInThe long version
How to Get Your Brand Cited by AI Search Engines
A playbook for earning citations in AI answers: how retrieval works, before and after passage rewrites, the page checklist, and how to track citation share.
Want this argument applied to your numbers?
No deck, no discovery sequence. Tell us what you are spending and where the pipeline stalls, and we will tell you what we would change first.
$9,137, credited back against a twelve month engagement.
Talk to an Expert