By Matija Konjić
- AI visibility needs its own measurement because assistant answers produce citations and referrals that standard analytics either hides or mislabels.
- Four inputs cover it today: referral traffic from assistant domains, systematic prompt sampling for mentions and citations, share-of-voice across your buying questions, and AI crawler activity in server logs.
- Assemble them into a one-page monthly scorecard, accept that attribution stays partial, and use the trend rather than the absolute numbers to steer the work that actually moves citations.
- Why AI visibility needs its own measurement
- The four inputs you can actually measure
- Input one: referral traffic from assistants
- Input two: systematic prompt sampling
- Input three: share of voice
- Input four: AI crawler activity
- Building the monthly scorecard
- The limits worth stating out loud
- Turning the numbers into work
Buyers increasingly ask an assistant before they ask a search engine, and the answer they get names two or three brands. If yours is absent from those answers, the loss never shows up as a ranking drop, a traffic dip or a red cell in any report you currently run. It shows up as pipeline that quietly stops arriving.
Measuring the surface is harder than measuring search and far from impossible. This guide covers the four things that can be tracked honestly today, how to assemble them into a scorecard a founder can read in a minute, the limits worth stating out loud, and how the numbers should change what the team actually does.
Why AI visibility needs its own measurement
The no-click problem
Classic search analytics assume a click. AI answers frequently resolve the question without one, Pew measured users clicking far less often when an AI summary appears, so the value moves from traffic to being named: the assistant recommends, the user acts later, and the credit lands on a branded search, a direct visit, or nothing traceable at all. A dashboard built for clicks reads that as flat performance while a competitor gets recommended all day.
There is also a discovery shift underneath. Assistants compress a results page into a handful of sources, which widens the gap between cited and uncited brands, exactly the dynamic behind getting cited in AI answers. Measurement matters because the surface rewards presence in a way you cannot infer from rankings.
A trend line for the board
There is a budget reason to measure early as well. Executives are already asking what AI means for the marketing plan, and teams answering with speculation lose the argument to teams answering with a trend line. Even a modest scorecard, built from free inputs, converts the topic from anxiety into a managed channel with a number attached.
The channel is young enough that measuring it at all is a competitive edge.
The four inputs you can actually measure
Everything trustworthy today comes from four sources.
None is complete alone. Referral data proves clicks happened but misses influence without clicks; sampling captures influence but only for the questions you thought to ask; crawler logs show access rather than outcomes. Read together they triangulate a picture accurate enough to steer decisions, which is the working standard for every emerging channel.
Start with two of the four if resources are thin: referral tracking, because it is a one-time setup, and prompt sampling, because it answers the question everyone is actually asking. Share of voice and log analysis can join in month two once the habit exists.
Input one: referral traffic from assistants
The GA4 segment
Set this up once and it runs. In GA4, build a segment or exploration filtered to session sources containing the assistant domains, chatgpt.com, perplexity.ai, copilot.microsoft.com, gemini.google.com, claude.ai, plus the newer entrants your niche shows. Track sessions, engagement and conversions separately from organic, because assistant referrals behave differently: fewer visits, longer sessions, higher intent, worth 4.4 times the value of an average organic visit by Semrush’s measurement, since the assistant already did the qualifying.
Cited landing pages
Watch the landing pages too. Assistant referrals cluster on the pages that got cited, which tells you which of your content the systems trust, and that list is usually more useful than the traffic total. Volumes stay modest for most sites today, so read the direction of travel rather than the size.
One configuration detail saves confusion later: some assistant traffic arrives with referrer stripped or bundled under direct, so expect undercounting and avoid comparing your assistant numbers with anyone else’s. Internal consistency month to month is what makes the line readable.
Input two: systematic prompt sampling
This is the input most teams skip and the one that answers the real question. Fix a list of ten to twenty prompts covering how buyers actually ask for what you sell:
- The category question. What the product class is and who it is for.
- The best-provider question. Who the assistant recommends when asked directly.
- The comparison question. You against the named competitor buyers shortlist.
- The how-do-I question. The task your product exists to solve.
- The price question. What things cost and what drives the range.
Run the same list across the major assistants on the same day each month, and record three things per answer: whether your brand appears, whether a link to your site is cited, and which competitors are named.
The fixed panel
Consistency is the whole method. Same prompts, same order, same day, logged in a sheet, because assistant answers vary between runs and only a fixed panel produces a trend rather than an anecdote. Half an hour a month buys the only direct evidence of whether the AI surface knows you exist.
Recording the wording
Log the exact answer text for your top five prompts rather than only the yes-or-no verdict. How you are described matters commercially: being named as the budget option, the enterprise choice or the specialist shapes who arrives, and correcting a wrong characterization is its own piece of work once you can see it.
Input three: share of voice
Turn the sampling log into a competitive number: across your question set, in what share of answers does each brand appear? Share of voice is more honest than a mention count because it normalises for how many brands each answer names, and it converts an abstract worry into a scoreboard leadership understands. Track your share, the leader’s share, and the gap.
The pattern to expect mirrors classic authority: the brands with the most external coverage and citations dominate the answers, which is why AI share of voice moves in step with the work described in brand search as a moat. When your share rises after a quarter of coverage and placements, that is the channel confirming the strategy.
Segment share of voice by question type too. Many brands hold decent presence on how-to questions and none on the buying questions that convert, and that split tells you exactly which content and coverage the next quarter needs.
Check robots and firewall rules for AI agents specifically as part of this input, since security tooling blocks unfamiliar bots by default and plenty of sites have been invisible to assistants for months without anyone choosing that.
Input four: AI crawler activity
Server logs show which AI crawlers fetch your pages and how often, and the list keeps growing: GPTBot, ClaudeBot, PerplexityBot, Google-Extended and others. Two things matter. First, access: confirm you are not blocking crawlers you want citing you, a check that belongs beside your llms.txt setup. Second, appetite: which sections get fetched repeatedly, since heavy crawling of a content cluster often precedes citations from it.
Treat logs as a leading indicator rather than a scoreboard. Fetching is not endorsement, but a site that never gets fetched cannot be quoted, and blocked or thin crawling explains absence faster than any theory about content quality.
Humans read the answers
Automate what deserves automating and no more. The referral segment and log parsing can run themselves; the prompt panel benefits from a human reading the answers, because the interesting findings are usually in the wording rather than the counts.
Building the monthly scorecard
One page, five lines, same format every month: assistant referral sessions and conversions; your mention rate across the prompt panel; your citation rate, where an actual link appeared; share of voice against the two nearest competitors; and a note on crawler access and volume. Add one line of interpretation, what changed and what the team did, and the scorecard doubles as a decision log.
Review it beside the link and PR numbers rather than in isolation. The correlation is the point: quarters with strong coverage and quality placements are the quarters citation rates climb, which is how the AI line justifies the budget that produced it.
Keep the raw sheets alongside the summary. Six months of logged answers becomes the only historical record of how the surface treated your category, and nobody else will have it.
Circulate it to whoever owns content and PR as well, since they are the people whose work the numbers actually reflect.
The limits worth stating out loud
Honesty protects the program, so the limits get stated plainly:
- Chosen blind spots. Sampling covers only the questions you thought to ask.
- Answer variance. Responses shift by user context, region and model version, so the panel measures a slice rather than the truth.
- Undercounted influence. The most valuable outcome, a recommendation acted on days later, arrives as branded search or direct.
- No vendor escape. The emerging AI-visibility platforms sample the same way you would, just at scale.
Error bars in the report
State those limits in the report itself. A scorecard that admits its error bars survives scrutiny and keeps its budget; one that implies precision gets dismantled the first time someone tests a number.
Expect the tooling landscape to churn as well, so keep your method portable: a spreadsheet, a fixed prompt list and a definition of each metric survive vendor changes that dashboards do not.
Turning the numbers into work
The scorecard earns its keep by directing effort. Low mention rate with healthy crawling means an authority problem: the systems can read you and choose others, which points at coverage, citations and expert visibility. Low crawling means an access or structure problem, fixable technically. Mentions without citations means you are known but not linkable on the topic, usually solved by publishing the referenceable material assistants prefer to point at. Competitor dominance in one question cluster names your next content and PR target precisely.
Every one of those responses is the same off-page work this blog keeps returning to, aimed by evidence instead of vibes. That is the honest state of AI measurement today: partial numbers, real signal, and a clear line from the scorecard to the authority building that moves it.
Set a review cadence of quarters for strategy and months for the numbers, matching the clock the underlying authority work runs on.
Then let the scorecard argue for the budget that improves it, month after month.
Want to know whether AI assistants are naming your brand or your competitor?