AI Visibility Metrics That Actually Matter (and Three That Don't)
By Keith Schilling · August 7, 2026 · 6 min read
New channel, new metrics, same old failure mode: dashboards fill with numbers that move impressively and mean nothing. AI visibility is young enough that the metric conventions are still being set, so here's the short list we've found actually supports decisions — and the seductive ones that don't.
The four that earn a slide
Mention rate — the share of category answers in which your brand appears at all. The foundation metric; the AI-era analogue of share of voice. Always report it per engine, because engines diverge enormously (we've measured 81% on one engine and 43% on another for the same brand set in the same month).
Recommendation rate — the share of answers where the engine actively suggests you, not just names you. This is the number that correlates with pipeline: "consider using X" moves buyers in a way passing references don't. The gap between your mention rate and recommendation rate is itself diagnostic — well-known-but-not-endorsed is a specific, fixable condition.
Citation share — how often your domain is used as a source, and which domains beat you. The companion list of *what engines cite in your category* is the most actionable artifact in all of this: it's your PR and content target sheet, ranked.
Consistency — of the times a prompt was asked, how often you appeared. Recommended in 3 of 3 runs is an owned position; 1 of 3 is a coin flip. Only measurable if your tracking runs each prompt multiple times, which is the quiet argument for repeat-run methodology.
A composite score, if you must have one number
Executives want one number, and a defensible composite exists: weight recommendation rate at 50%, mention rate at 30%, citation share at 20%, aggregated across your prompt set. That's the Treyci visibility score. The weights are judgment calls — what matters is that the formula is public, stable, and always shown next to competitors' scores computed identically. A score of 46.8 tells you nothing; 46.8 against your archrival's 43.3 tells you the race is close and worth running.
Three metrics to politely decline
Daily single-run deltas. Answers vary between askings; a daily chart of one-run samples is a random walk that will absorb hours of your team's storytelling. Blended cross-engine averages as a headline — the average of 81% and 43% describes no engine a buyer actually uses; keep engines separate. Raw answer counts ("we analyzed 12,000 responses!") — that's a receipt for compute spend, not a brand outcome. Volume is the denominator, never the result.
The test for any AI visibility metric is the same as for any metric: does a specific person change a specific behavior when it moves? The four above pass. Report those, trend them monthly, and let the rest stay in the export file.
Treyci tracks your brand across ChatGPT, Perplexity, Gemini, and Grok — every prompt run three times, every month, from $99.
See plansHow we measure →