I got asked last month to review a $2,400 a month AI visibility dashboard a client had already bought. It reported a 'brand visibility score' of 34, up from 29 the previous week. Nobody on the call could tell me what prompt generated that number, how many times it ran, or what an increase of five points was supposed to mean for pipeline. That is not measurement. That is a green arrow with a subscription attached to it.
You do not need to buy that. You need a prompt list, a spreadsheet, some referrer logic in your analytics, and the discipline to run it on a schedule instead of once. This post is the version of that setup I actually build before I let a client spend a dollar reacting to an AI visibility number.
Start with the prompt set, not the tool
The single most common mistake in AI search tracking is building a prompt list around your own brand name. 'What does [company] do' is a prompt that tests whether the model has heard of you. It is not a prompt that tests whether you show up when a buyer who has never heard of you is trying to solve a problem. Those are different questions and only the second one moves pipeline.
Build your prompt set the way you would build a keyword list for paid search, except the unit is a question, not a term. Pull from the actual language your sales team hears on discovery calls: the comparison questions, the 'best tool for X team size' questions, the 'alternative to [category leader]' questions, the 'how do I solve Y problem' questions that come before anyone has named a vendor. Thirty to fifty of these, refreshed quarterly as your positioning shifts, is enough to be statistically useful without becoming a full time job.
Why a single run is noise, not a number
Language models are not deterministic, even with the temperature turned down, and the retrieval layer sitting in front of them refreshes its index on a schedule you do not control. Ask the same buying question twice in the same hour and you can get two different citation sets. Ask it once a month and report the result as your visibility trend, and you are reporting the weather on one day as the climate.
You do not need 21,000 prompts. You need enough repeated runs of the same fixed set that a single odd answer does not swing your headline number. Run each prompt in the set five times per measurement period, on the same day where possible, and report citation rate as a percentage of runs rather than a binary yes or no. A page that got cited in three of five runs is meaningfully different from one cited in one of five, and neither is the same as being cited in zero.
The three metrics actually worth tracking
Everything else is decoration. If you track only these three, you will know more than most teams running a paid AI visibility platform.
- 1
Citation share
Out of every prompt run in your set where the model names at least one external source, what percentage of the time is one of those sources you. This is the closest equivalent to share of voice and it is the number to chart week over week, per assistant, since ChatGPT, Perplexity and AI Overviews cite differently and moving one does not move the others.
- 2
Mention position and sentiment
When you are cited, are you the first source named or the fourth. Are you described as a leading option, a budget option, or mentioned in passing as one of several. A citation buried fourth and hedged with 'some users report' is not the same win as being named first with a specific claim attributed to you. Log this by hand for now. No tool does it reliably yet.
- 3
Referral sessions and their conversion
How many sessions actually land on your site from an AI answer engine, and what do those sessions do once they arrive. Citation without a click is a brand impression, not pipeline. This is the metric finance will actually ask about, and it is the one most AI visibility dashboards do not attempt because it requires your own analytics data, not just prompt scraping.
Catching the traffic that Google Analytics misses
ChatGPT and Perplexity outbound clicks generally carry a clean referrer header. Filter your analytics for a referrer of chatgpt.com, perplexity.ai, and gemini.google.com and you will see genuine AI-sourced sessions, assuming your analytics platform has not already bucketed them under 'referral' or 'other' by default. Check that bucket. Most teams have been sitting on this data unlabeled for a year.
Google AI Overviews is the harder case. A click from inside an AI Overview panel frequently arrives with the same google.com referrer as a plain organic click, and on mobile apps it can arrive with no referrer at all. You cannot cleanly separate this from server logs alone using referrer data. What you can do is look at server log patterns for a shift in landing page distribution that does not match your organic keyword rankings, and cross reference against Search Console's 'AI Overviews' appearance data where it is available for your property. It is imperfect. Say so in your reporting rather than pretending the number is exact.
| Measurement approach | What it actually tells you | What it misses |
|---|---|---|
| Paid AI visibility dashboard | Mention count across a vendor-chosen prompt list, usually brand-heavy | Real buying-question performance, referral conversion, why you lost a citation |
| Referrer filtering in GA4/analytics | Confirmed sessions from ChatGPT and Perplexity with clean referrers | AI Overviews traffic disguised as organic, app-based referrer stripping |
| Server log and user agent analysis | Bot crawl patterns from GPTBot, PerplexityBot, and retrieval crawlers | Whether a crawl visit ever turned into a citation or a session |
| Manual prompt set run weekly | Citation share, position, and sentiment on questions that mirror real buyers | Scale; hand logging caps out around 50 prompts before it eats a full day |
| Pipeline-matched referral tracking | Whether AI-sourced sessions convert at a comparable rate to organic | Anything upstream of the click; tells you nothing about why you were or were not cited |
Building the tracking loop
The sequence below is how a team starting from zero can stand the loop up. It takes about one working day to build and roughly two hours a week to maintain once it is running.
- 1
Build the prompt set
Pull 30 to 50 real buying questions from sales call transcripts and support tickets, not from a keyword tool. Weight it toward comparison and problem-first questions, cap brand-name prompts at 20 percent of the list.
- 2
Establish the baseline
Run the full set five times each against ChatGPT, Perplexity, and Google AI Overviews before you change anything. Log citation yes or no, position, and one line describing sentiment for every run in a shared sheet.
- 3
Instrument referral capture
Add referrer segments for chatgpt.com, perplexity.ai, and gemini.google.com in your analytics tool. Flag AI Overviews as a separate hypothesis to test via Search Console rather than a confirmed number.
- 4
Run on a fixed weekly cadence
Same day, same prompts, same models, five runs each. Anything less than weekly and you will not catch the effect of a content change within a useful reporting window.
- 5
Tie movement to changes
Keep a one line changelog of every content or schema change you ship, dated. When citation share moves, check the changelog first before assuming the model got smarter or dumber on its own.
Tying visibility movement to content changes
This is the step almost everyone skips, and it is the reason so many AI visibility reports read like weather reports instead of causal arguments. A citation share number that moves from 12 percent to 22 percent means nothing on its own. It means something when you can point at the dated content change that happened two weeks earlier and show the timing lines up.
Worked example
Illustrative modelA pricing page rewrite and a four week lag
An illustrative model, not a client account. A pricing comparison page is rewritten in early June to replace adjective-led claims with specific, sourced numbers: an exact onboarding timeline, a support response time in hours, and a dated benchmark statistic. The prompt set included six comparison-style buying questions relevant to that page.
- Citation share on those 6 prompts, May baseline (5 runs each)
- 13 percent
- Citation share, week of the rewrite
- 17 percent
- Citation share, 4 weeks after the rewrite
- 41 percent
- Referral sessions from ChatGPT and Perplexity to that page, May
- 9
- Referral sessions from ChatGPT and Perplexity to that page, week 8 post-rewrite
- 31
Result: The lift did not show up the week of the change. It showed up over the following month as the retrieval layer re-indexed the page and the model started treating the new, sourced claims as citable. Reporting the number the week of the change would have looked like the rewrite failed. Reporting it a month later showed a real threefold increase in citation share on exactly the prompts the page was meant to win.
The strongest case against this
This is a lot of manual work for a channel that still drives a small fraction of total traffic. Wouldn't the hour be better spent on paid or organic where the volume is already there and the tracking is solved?
That's a fair allocation argument today and I would not tell a resource constrained team to build a full weekly tracking loop before they have their paid and organic reporting in order. But the volume argument has a shelf life. AI referral sessions are still a minority of B2B traffic, and the teams building the measurement discipline now are the ones who will have two years of trend data when it stops being a minority. The actual cost here is not high. It's one working day to set up and two hours a week to maintain. Compare that to what most teams already spend arguing about attribution models for channels that have been tracked for a decade, and the AI visibility loop is cheap by comparison. Skip the paid dashboard, not the discipline.
What tools genuinely save you time on
I am not against paying for tooling. Running 50 prompts five times a week by hand across three models is tedious and a scraper that automates the API calls and logs the raw output for you is worth money. What is not worth money is a vendor's proprietary 'visibility score' that compresses citation rate, mention count, and sentiment into one number you cannot decompose or audit. Buy the automation. Do not buy the black box score. Keep your own prompt list, your own logging, and your own definition of what counts as a citation, and use the tool to save you the API calls, not to replace your judgment about what the numbers mean.
The other thing tools reliably get wrong is attributing a lost citation to the wrong cause. A tool will tell you a competitor took a citation you used to hold. It will not tell you whether that happened because their page added a sourced statistic, because a G2 review update changed the consensus description of your category, or because the model's retrieval index simply refreshed and picked up a newer page. You have to read the actual output and compare it to what changed. That comparison is the job. The scraping is the plumbing.
What a good citation actually looks like in the raw output
Most teams never read the raw model output. They read a vendor dashboard that has already decided what counts as a mention. Go read the actual text a few times and you will notice citations are not binary. A model can name you as one of six options in a numbered list with no elaboration, or it can pull a specific claim from your page and attribute it directly, sometimes with a link and sometimes without. Those are wildly different outcomes for pipeline even though a naive tracker logs both as 'cited: yes.'
Sort every citation in the raw outputs into three buckets: named in a list with no detail, named with a generic descriptor lifted from a category page, and named with a specific attributed claim. Only the third bucket gives a reader a reason to leave the chat window, so it is the only one with a plausible path to a click. If your tracking only counts appearances, you are optimizing for the first bucket, which is the cheapest and least useful one to win.
Sentiment is not a nice-to-have column
A mention with a qualifier attached, 'some users report', 'may be a good fit for smaller teams', 'less established than', is doing different work than a mention stated as fact. Score sentiment on a simple three point scale: neutral factual, positively framed, or hedged and negatively framed. Track it separately from citation share. Citation share can climb while sentiment gets worse, because competitor review content gets folded into how the model describes category weaknesses. A rising citation count on its own would read as a win.
Sampling cadence in practice, not in theory
Weekly is the floor, not the ideal. In practice I run the full prompt set on a fixed day, usually Monday morning before anyone on the team has made a content change that day, so the previous week's shipped work has had time to be crawled and possibly re-indexed. Running immediately after a content push tells you nothing, because most retrieval layers have not refreshed yet. Running four weeks later tells you whether the change actually mattered, once you have a comparison point.
I also keep a rolling four week average alongside the weekly number, because a single week can wobble ten points in either direction even with five runs per prompt. A CMO does not need to see week over week noise. They need to see whether the four week trend line is moving, and in which direction, on the buying-question prompts that matter to pipeline rather than the ones that flatter the brand.
Reading server logs without a specialist tool
You do not need an enterprise log analytics platform to catch AI crawler activity. Most servers keep raw access logs, and GPTBot, PerplexityBot, ClaudeBot, and Google-Extended all identify themselves in the user agent string. Grep your access logs for those strings over a 30 day window and you will see which pages are actually being crawled for retrieval, which is a leading indicator that predates any citation showing up in your prompt set.
This matters because a page that never gets crawled by these agents cannot be cited, no matter how well it is written. If your access logs show zero hits from GPTBot on your highest priority page after 60 days, you have a crawl access problem, likely a robots.txt block or a JavaScript rendering issue hiding content from a bot that does not execute scripts, and no amount of prompt engineering or content rewriting will fix that until the crawl access is resolved first.
What to actually put in a monthly report
A CMO does not want to see 50 rows of raw prompt data. They want four numbers and one sentence of context per number. Citation share this month versus last, broken out by assistant since they move independently. Average mention position when cited. Referral sessions from AI engines and their conversion rate compared to organic. And one line connecting any meaningful move to a specific, dated content or outreach change. If you cannot fill in that last line, say so plainly rather than inventing a causal story to make the number look intentional.
- Citation share by assistant, this period versus last, on the fixed buying-question prompt set
- Average mention position and sentiment score when cited, not just a yes or no
- AI referral sessions and their conversion rate against your organic baseline
- One dated change tied to the largest single move in either direction
- Crawl access status for any page that should be getting cited but is not showing up in logs
The failure mode on the other side
The opposite mistake is just as common as buying the vanity dashboard: tracking nothing at all because the channel still feels experimental. I understand the instinct. AI referral sessions are a fraction of organic traffic for most B2B sites today, and it is tempting to treat the whole category as a rounding error not worth instrumenting. The problem with that stance is that the retrieval layers making citation decisions today are the same ones that will be making them at scale in two years, and the teams with two years of clean prompt-run history will be the only ones able to tell a credible story about what actually moved their visibility versus what the model changed on its own.
Set up the tracking now, run it lean, and resist the urge to make it bigger than it needs to be. A spreadsheet, 30 prompts, five runs a week, and a changelog. That is the whole system. Everything past that is optional until the channel earns a bigger budget on its own numbers.
The close
None of this is complicated. It is unglamorous, it takes a spreadsheet and a recurring calendar block, and it will never produce a dashboard slick enough to put in front of a board without context. What it will do is tell you, with actual evidence, whether the content and outreach work you are funding is moving the number that matters. Everything else being sold as AI search measurement right now is a mention counter wearing a strategy's clothes.
What I would do Monday
- 1Write 30 real buying-question prompts from your category, none of them containing your brand name.
- 2Run the set against ChatGPT, Perplexity and Google AI Overviews five times each and log citation, position and sentiment per run.
- 3Add referrer and user agent filters to your analytics for chatgpt.com, perplexity.ai, and AI Overviews click patterns.
- 4Pull the last two content changes you shipped and check whether citation share moved on the pages they touched.
- 5Kill any tracking dashboard in your stack that reports a mention count with no prompt list attached to it.
Common questions
- What is the best way to track AI search visibility?
- Build a fixed set of 30 to 50 prompts that mirror real buying questions in your category, run them against the same models on a weekly cadence, and log three things: whether you are cited, where in the answer, and what tone the mention carries. Pair that with referrer detection in your web analytics and server logs for ChatGPT, Perplexity and Google AI Overviews traffic, then tie both to pipeline. A tool can automate the running and logging. It cannot decide what a real buying question looks like for your category, and that decision is most of the work.
- How often should I check my AI citation rate?
- Weekly at minimum, monthly at the absolute floor if you are resource constrained. Model outputs are not deterministic even at low temperature, and retrieval layers refresh their index on their own schedule, so a single run on a single day tells you what happened on that day, not what is generally true. Run the same prompt set at least five times per period and report the citation rate as a percentage of runs, not a yes or no.
- Can Google Analytics track ChatGPT and Perplexity referral traffic?
- Partially. Standard analytics tools will show you sessions with a referrer of chatgpt.com or perplexity.ai when the click carries a referrer header, which most AI answer engine outbound links do. What GA4 will not show you is traffic from AI Overviews inside Google search, because that traffic usually arrives with a google.com referrer identical to organic search, or none at all on mobile apps. You need server log analysis and user agent detection to separate that traffic from regular organic.
- Do AI visibility tracking tools tell you why you are not being cited?
- No, and this is the gap that gets sold over. Most tools report a mention count or a citation percentage per prompt. None of them explain that you lost the citation because your page made an unsourced claim, because a competitor's G2 profile was more specific, or because your content was published after the model's last retrieval pass. You still have to do the manual diagnostic work of reading the actual model output and comparing it to what got cited instead.
- What counts as a real buying-question prompt versus a vanity prompt?
- A vanity prompt asks the model about your brand directly: 'what is [company] known for.' A real buying-question prompt is what a prospect actually types before they have heard of you: 'best tools for tracking sales pipeline forecast accuracy' or 'alternatives to [category leader] for a 50 person sales team.' Vanity prompts tell you whether the model has heard of you. Buying-question prompts tell you whether you show up when it matters, which is the only number worth reporting to a CMO.
Where these numbers come from
Download citations (JSON)Each claim below names its source and how recent that source is. Anything marked as a model is an illustration with stated assumptions, not measured market data.
- 21,075CurrentCurrentA AI search behaviour is treated as usable for 3 months. This one is comfortably inside that window, and is re-checked before 2026-10-13. Use the figure as stated.SourcedSourcedA named, dated third-party publication backs this number. The source, its publisher and its publication date are listed below the claim.
AI engine responses and 25,337 citations were tracked across ChatGPT, Perplexity, Gemini, Google AI Overviews and AI Mode over a three month window to produce a…
AI Search Citations Study: What 25,000+ Citations Reveal — DeltaV Digital, July 2026
- 25,337 citationsCurrentCurrentA AI search behaviour is treated as usable for 3 months. This one is comfortably inside that window, and is re-checked before 2026-10-13. Use the figure as stated.SourcedSourcedA named, dated third-party publication backs this number. The source, its publisher and its publication date are listed below the claim.
Citations tracked across ChatGPT, Perplexity, Gemini, Google AI Overviews and AI Mode in a single study, which is the scale you are sampling against when you ha…
AI Search Citations Study: What 25,000+ Citations Reveal — DeltaV Digital, July 2026
- Referrer filtering in GA4/analytics / Scale; hand logging caps out around 50 prompts before it eats a full dayIllustrative modelIllustrative modelThis is an illustrative model with stated assumptions, not measured market data. Treat it as arithmetic you can re-run with your own inputs, never as a benchmark.No external studyNo external studyNo external study is attached to this figure. It is either an internal illustration or a number describing the shape of an argument rather than a market measurement.
No single row is sufficient on its own. The combination is the actual tracking system.
Last updated
- First published
- Last updated
- Last fact review
Who wrote this
Avishai Sam Bitton
Founder, DemandBox
Avishai runs demand generation programs for B2B SaaS companies across performance marketing, SEO, and answer engine optimization. He works directly with the teams he advises, with no account managers in between.
Connect on LinkedInThe long version
How to Get Your Brand Cited by AI Search Engines
A playbook for earning citations in AI answers: how retrieval works, before and after passage rewrites, the page checklist, and how to track citation share.
Want this argument applied to your numbers?
No deck, no discovery sequence. Tell us what you are spending and where the pipeline stalls, and we will tell you what we would change first.
$9,137, credited back against a twelve month engagement.
Talk to an Expert