The metrics that actually predict growth

Reporting

The metrics that actually predict growth

Filed under

Reporting

Published

Vanity numbers feel good and tell you nothing. Here are the signals we watch when we want to know if the work is really working.

You can measure whether AI engines mention your business, cite your content and prefer your competitors. What you can't do is measure it precisely, because the thing being measured doesn't hold still - the same question, asked twice, returns different answers. Credible measurement in AI search means a fixed panel of buyer questions, sampled repeatedly across engines, with mentions and citations tracked separately, and progress judged on trends rather than snapshots. Everything else is decoration.

That's the honest position, and it matters because a measurement category this young attracts a lot of confident dashboards. One agency group put it bluntly in its own published guidance: "none of your AI visibility data is accurate" - the mention rates fluctuate run to run. The answer isn't to give up on measurement. It's to measure like you mean it.

The metrics, and what each one actually tells you

Mention rate is the share of relevant answers that name your brand, linked or not. It tracks brand trust and positioning. Citation rate is the share of answers that credit your content as a source. It tracks topic authority and content quality. They are not the same thing: Semrush's ghost citations research found 62% of AI citations come with no brand mention at all, and the engines lean opposite ways - Gemini names brands far more often than it links them, while ChatGPT links far more often than it names. Report the two separately or you will fix the wrong problem.

Share of AI voice is your mentions (or citations) as a share of the total across a defined set of competitors, on a fixed set of prompts. It's the metric that gives the others context, and it's meaningless without a named competitor set. Position tracks how often you're the first brand raised. Sentiment tracks how you're framed when you appear. Prompt coverage tracks how many of your buyers' questions surface you at all. And referral plus conversion tracks what happens when someone actually clicks through.

None of these comes from logged impressions the way Google Search Console data does. Every number is an estimate built from sampling. That's the defining limitation of the category, and any report that doesn't say so is overclaiming.

Why the numbers wobble

AI answers are non-deterministic: the same prompt produces different results across runs. They're personalised by account, history and location. And they vary by industry - one study of 177 brands across eight platforms found legal firms are cited constantly but rarely named, while financial services skew the other way, with editorial media getting the citations. So a week-to-week drop in your mention rate is usually noise. The signal is the trend over eight to twelve weeks against the same prompt panel.

A method you can defend

Five steps, none optional.

Build a fixed prompt panel from real buyer questions - the many ways people actually ask about your category, not the phrasing you wish they used. Sample it repeatedly across the engines that matter to your market, on a set cadence. Lock the counting rules before you start: does a passing mention in paragraph three count the same as a first-line recommendation? Definition drift ruins more AI measurement programs than the models do. Track deltas, not absolutes - report movement against your own baseline and your named competitors. And publish the methodology inside the report, because in a field this young, how you counted is part of the finding.

The tools, and what they don't do

The AI visibility category raised more than US$300 million in funding between mid-2025 and early 2026, so there's no shortage of options. As at mid-2026, prices in US dollars, all directional:

Otterly starts around US$29 a month and is the most accessible entry point. Peec AI, from about US$95 a month, is the strongest mid-market and agency option. Semrush's AI toolkit runs about US$99 per domain and bolts onto an existing suite. HubSpot's AEO Grader is US$50 a month standalone. Ahrefs Brand Radar builds its prompts from real search data rather than synthetic queries and keeps historical records, but realistically lands above US$800 a month all-in. Profound is the enterprise leader - customers include Target and Walmart - with deployments commonly in the US$2,000 to US$5,000 a month range and above.

Two things every tool has in common. They sample synthetic prompt panels, not real buyer query logs, so they estimate rather than observe. And they measure - none of them fixes anything. A dashboard is a thermometer, not a treatment.

What AI referrals look like in your analytics

Small, hidden, and unusually valuable. Across large multi-site studies, AI referrals sit around 1% of sessions for typical business sites. GA4 buries them by default - visits from ChatGPT, Perplexity and Claude land in Referral, Unassigned or Direct - so set up a custom channel group matching chatgpt, perplexity, claude, gemini and copilot as sources. Google has shipped a native AI channel in GA4, but it misses roughly 30% of sources, including Perplexity, so run both. And traffic from Google's own AI Overviews and AI Mode is effectively invisible, because Google doesn't attribute it separately.

The value side is directional but consistent. Ahrefs reported that AI assistants drove roughly half a percent of its visitors but 12% of its signups. Adobe's Q1 2026 data had AI-referred shoppers converting 42% better than average and spending 48% longer on site. The plausible mechanism: the AI has already done the comparing, so the visitor arrives late-stage and pre-qualified. Two caveats belong in any report that quotes these numbers: the most-cited per-platform conversion figures trace back to a single B2B case study, and the absolute volumes remain small - ChatGPT handles roughly 12% of Google's query volume but sends about 190 times less traffic to websites.

What a credible report looks like

A per-engine breakdown, never one blended score. Mention rate and citation rate as separate lines. Share of voice against named competitors. Trend over at least eight weeks, not a snapshot. Prompt coverage. Referral and conversion from a properly configured GA4 channel. And a plain-language sentence against every number explaining why it matters, because a metric without a consequence is furniture.

What to refuse: single-run screenshots, absolute scores with no competitor context, and any report that won't show you its prompt panel and counting rules.

What moves these numbers in the first place is covered in what kind of content AI actually cites (/journal/what-content-ai-actually-cites).

FAQ

What's a good mention rate? No credible external benchmark exists yet. Judge yourself against your named competitors and your own trend line.

Which tool should we buy? It depends on stage and budget. Entry tools from about US$29 a month are enough to establish a baseline. Whatever you choose, remember it measures - the fixing is content and brand work.

Why did our score drop this week? Probably noise. AI answers vary between runs. Judge the trend over eight to twelve weeks.

Can we see Google AI Overviews traffic in GA4? Not directly. Google doesn't attribute it separately as of mid-2026.

Sources

  • Semrush - The ghost citations study (2026): semrush.com/blog/the-ghost-citations-study

  • Brainlabs - published guidance on the accuracy of AI visibility data (2026)

  • Victorious - AI citations and mentions across 177 brands and eight platforms (2026)

  • Conductor - AI citation and traffic research (2025-26): conductor.com/academy/how-ai-citations-differ

  • Ahrefs - AI assistant referral and conversion data; Brand Radar methodology (2025-26): ahrefs.com

  • Adobe - Q1 2026 digital insights on AI-referred shoppers

  • Kumar et al. - non-determinism and citation behaviour in AI engines (2025): arxiv.org/abs/2509.10762

  • Vendor pricing pages for Otterly, Peec AI, Semrush, HubSpot, Ahrefs and Profound, mid-2026. All figures hedged as directional.

What kind of content does AI actually cite?

The best CMS for AI search: WordPress vs Webflow vs Squarespace vs Framer

The best CMS for AI search: WordPress vs Webflow vs Squarespace vs Framer

Get In Touch

Let’s build something people remember

Talk to strategists

No account managers - you speak directly with the people running the marketing.

hello@netradigital.com.au

(61) 468 939 110

Melbourne & Berlin

Get In Touch

Let’s build something people remember

Talk to strategists

No account managers - you speak directly with the people running the marketing.

hello@netradigital.com.au

(61) 468 939 110

Melbourne & Berlin

Get In Touch

Let’s build something people remember

Talk to strategists

No account managers - you speak directly with the people running the marketing.

hello@netradigital.com.au

(61) 468 939 110

Melbourne & Berlin