Visibility Overview
Fires real category prompts at ChatGPT, Gemini, Perplexity, Claude and Google AI Overview, then scores how often your brand is named, where, and in what words.
Last updated 2026-09-14
Summary#
Visibility Overview measures whether AI assistants name your brand. It generates ten real category questions that never mention you, sends them to every configured AI engine, and reports how often you appeared, at what position, with what sentiment, and which sources the engines cited. Its internal tool id is ai_visibility.
Purpose#
A growing share of buying research now happens inside an AI assistant. Someone asks "what's the best tool for X", gets three names and a paragraph each, and never opens a search results page. If you are not one of those names, you were never considered — and nothing in your analytics will tell you it happened.
This tool exists to make that invisible surface measurable. The decision it supports is whether AI visibility is a problem worth investing in for your brand, and if so, on which engine and against which competitor.
Overview#
You enter a brand or domain. Metric Vault first writes the questions: it grounds them in your domain's own top-ranking keywords and recent news coverage, then produces a mix of category-recommendation, comparison, problem-solving and pricing questions. Crucially, none of them contain your brand name. If you come back in the answer, you earned it.
Those ten questions go to every AI engine that is configured on the platform — ChatGPT, Gemini, Perplexity, Claude and Google's AI Overview. Each answer is scanned for your brand: whether it appears, how early in the answer it appears, whether the surrounding language is positive, neutral or negative, and which sources the engine cited.
The headline figure is simple arithmetic: appearances divided by the number of chances, where chances is questions multiplied by engines. Ten questions across five engines is fifty chances.
The run is genuinely live. It queries the models each time, which is why two runs of the same brand can differ slightly.
Benefits#
- A measured number for AI visibility rather than an assumption.
- Per-engine breakdown, so you know which assistant to work on.
- Verbatim excerpts — the actual sentences the models produced about you.
- A citation-source list, which doubles as a placement target list.
- Feeds Growth Actions, so a run here unlocks a ranked action queue.
- Costs 1 credit, which keeps it under the monthly premium threshold.
Use Cases#
Establishing a baseline. Before investing in AI visibility, you need the number you are trying to move.
Diagnosing a weak engine. Strong on ChatGPT and near zero on Perplexity is a specific, fixable problem with a specific fix.
Finding content gaps. Every question where you did not appear is a brief.
Building a placement list. The citation sources are the domains the engines already trust in your category.
Checking how you are described. The excerpts show the exact words buyers see, which no score can convey.
Requirements#
- A signed-in account.
- A paid plan — 1 credit per run, and any cost above zero is blocked on Free.
- A brand with some public footprint. A brand-new name returns an honest empty result rather than an invented score.
- At least one AI engine configured on the platform. Which engines are live is a platform setting, not an account setting — see AI models used across the platform.
- Available credits and hourly headroom.
Permissions#
| Your situation | What you see |
|---|---|
| Signed out | Please sign in to run this tool. |
| Free plan | This tool needs a paid plan. Free includes the 10 technical SEO tools; upgrade to Pro to unlock the rest. |
| Paid plan | Full access |
| Pro and above | The Get-recommendations panel under the result |
| Any team role | Identical. Access is by plan, not by role |
Cost#
1 credit per run through the dashboard. The run button reads Analyze AI presence · 1 report.
Below the premium threshold of 3, so the monthly quota never blocks it — only the hourly fair-use cap of 100 light-tool calls per hour. A shared-cache hit still charges the credit; the cached window for this tool is two days, so a re-run inside that window can return the earlier measurement. Use the free public AI-visibility checker if you only want a technical readiness score. See How credits work and Result caching and freshness.
Navigation Path#
Dashboard → AI Visibility → Visibility Overview
The in-view title is AI Visibility Score.
Inputs#
| Field | Accepts | Required | Default | Validation | Notes |
|---|---|---|---|---|---|
e.g. nike.com | A brand domain, or a brand name | Yes | Empty | Empty input shows Enter a value first. | A domain works best: the prompt generator grounds itself on that domain's ranked keywords and news coverage |
| Try sample | — | — | — | — | Runs with whatever is in the field and charges a credit |
TRY: chips (nike, tesla.com, salesforce.com, figma.com) | — | — | — | — | Clicking one fills the field and runs immediately, charging a credit |
| Country picker (topbar) | Country list | No | United States | — | Sent with the run |
Pressing Enter in the field runs the tool.
Step-by-Step Guide#
- Open
AI Visibilityand choose Visibility Overview. - Enter your brand domain in the field marked
e.g. nike.com. - Click Analyze AI presence · 1 report, or press
Enter. - Wait for
Analyzing… this can take a few seconds for live data.to clear. This run queries live models and is slower than a data lookup. - Read the five KPI tiles, then the
Sample Verbatim LLM Excerptspanel. - Work down through the per-engine breakdown, the action plan and the citation sources.
- Export or share the report from the buttons at the top of the view.
Reading the Results#
The header. Your domain, an AI Visibility badge and a Live LLM tracking badge, then a line stating exactly what was measured: "Real measurements from 10 category prompts × N AI engines … Brand surfaced in X of Y response slots." Read that sentence before anything else — it is the denominator for everything below.
The five KPI tiles.
| Tile | Reads | Good | Bad | Action |
|---|---|---|---|---|
Visibility | The score out of 100, with a band | 60+ | Under 30 | The band is the summary; see the table below |
Appearances | X/Y with the prompts × engines breakdown | Above half | Under a fifth | This is the raw evidence behind the score |
Sentiment | Positive, Neutral, Negative or No mentions, with the mention count | Positive | Negative | Negative with real mention volume is more urgent than a low score |
Avg position | #N — how early you appear when you do | #1 or #2 | #6 and beyond | Appearing last in a list of eight is close to not appearing |
Engines | N/5 with the engine names | 5/5 | 1/5 | A low number means engines are not configured, not that you scored badly |
The bands. The word next to the score is what you report upward:
| Score | Band | What it means | What to do |
|---|---|---|---|
| 80-100 | Dominant | You are the default answer in your category | Defend it; track it weekly |
| 60-79 | Healthy | Named in most relevant answers | Push on the engines and prompts where you are absent |
| 30-59 | Emerging | You appear when the question is close to your niche and vanish when it widens | Build citable, factual content on the adjacent topics |
| 10-29 | Limited | Occasional, incidental appearances | Treat AI visibility as a project, not a tweak |
| 0-9 | Invisible | The assistants effectively do not know you in this category | Start with third-party presence, not more marketing copy |
A score of 0 is reported plainly: the brand was not surfaced by any tested prompt, and because the prompts are real category buyer queries, that is real AI invisibility rather than a measurement failure.
Sample Verbatim LLM Excerpts. Read this before the charts. Three quote cards — a Brand surfaced answer, a Mixed coverage answer, and a Competitors won answer — each showing the engine, the exact prompt, and the literal words the model returned. The badge reads REAL · LIVE LLM.
The losing card is the most useful thing on the screen. It shows which competitors got named instead of you and what the model said about them. That is your content brief.
AI Visibility Score. A Composite score meter repeating the headline, with the threshold stated: 80+ is dominant in your category.
Per-Engine Visibility — appearance rate. A bar per engine showing what share of the ten prompts surfaced you there. Read the differences, not the average:
- Strong on ChatGPT, weak on Perplexity usually means you have general brand recognition but few current, citable web sources. Perplexity leans on live retrieval. Fix it with well-structured factual pages and third-party coverage, not more brand copy.
- Weak everywhere except one engine means the one engine has a source the others do not — find it in the citation list.
- Zero on Google AI Overview is the highest-traffic blind spot, and the recommendations panel flags it as such.
Recommended Action Plan — prioritized. Concrete steps grounded in this analysis. When an engine is not configured, the first step is to add its key; when citation sources exist, one step names the specific domains to pitch.
What This Measurement Tells Us. The written findings. One always states the run size and the score. Another names any engine that was not measured because no key is configured — that is why a missing engine row is an explained absence, not a zero. When Google AI Overview ran, a finding states for how many of the ten prompts an AI Overview appeared at all and in how many of those you were named; Google only shows an AI Overview for some queries, so that count is itself a signal.
Visibility Trend — across 5 AI platforms. Your measured score over time, built from stored runs. A single run has almost no trend; the value appears after several. Because each run generates its own questions and models are non-deterministic, judge the trend line rather than any single reading.
Representative Prompts. The questions that were asked and which engines named you for each. The ones where nothing named you are the opportunity list.
Top Citation Sources. The domains the engines drew on, with an observed citation count. This is the single most actionable list in the report: those are the pages you want to be present on. Perplexity cites most heavily, so this list is largely what it trusts in your category.
Sentiment by Platform. Positive, neutral and negative counts per engine, over the answers where you were named. An engine that describes you worse than the others usually traces back to one dominant source — find it in the citation list and address it there.
Get recommendations. A panel under the result converts the findings into an ordered plan, grouped as Fix first, Do next and Quick wins: a low overall score, negative sentiment, absence from Google AI Overview, a competitor outranking you, and any engine with zero appearances. It is a Pro-and-above feature with its own monthly allowance — see Get Recommendations.
What to do with it. Take the losing excerpt and write the page that would have won that answer. Take the top three citation sources and pitch them. Re-run monthly and watch the trend, not the single number.
Examples#
Example: You run ourbrand.com. The header states 10 prompts across 4 engines and the brand surfaced in 11 of 40 slots. Visibility reads 28/100 with the band Emerging; Appearances reads 11/40; Sentiment reads Positive on 11 mentions; Avg position is #3; Engines reads 4/5. Per-engine, ChatGPT is 45% and Perplexity is 10%. The Competitors won excerpt shows Perplexity recommending two rivals for "best X for small teams" and citing a review site you have never been covered on. Top Citation Sources lists that review site first. Your plan writes itself: get covered on the review site, publish a factual comparison page for that exact question, and re-run in four weeks.
Screenshots#
Tips#
- Use the domain rather than a bare brand word. The prompt generator grounds itself on the domain's ranked keywords, and a common word inflates false matches.
- Read the excerpts before the score. The words tell you more than the percentage.
- Save the citation-source list. It is your placement target list for the quarter.
- Run it on a competitor to see the same five measures for them.
- Monthly is the right cadence. Re-running the same day inside the two-day cache usually returns the same figures for the same credit.
Best Practices#
- Track the score over several runs rather than reacting to one. Each run writes its own questions and the models are non-deterministic.
- Turn every prompt where you were absent into a content brief, using Topic Map or Content Brief & Template.
- Run this before Growth Actions. It is the richest source the growth queue reads.
- For continuous coverage of specific questions, move to Prompt Tracking and the AI-visibility alerting described in AI prompt tracking.
Common Mistakes#
- Reading a missing engine as a zero. If an engine has no key configured on the platform, it did not answer; the findings name it.
- Comparing scores across brands run on different days without noting that each run generates its own question set.
- Assuming the tool asked the models about you by name. It deliberately does not — that is what makes an appearance meaningful.
- Judging a brand-new company harshly. An
Invisibleband on a six-month-old brand is expected, not alarming. - Ignoring sentiment because the score is fine. Being named negatively is worse than not being named.
Limitations#
- Ten generated questions per run. It is a well-designed sample, not an exhaustive audit of everything an assistant might say.
- The engines queried depend on which providers are configured on the platform.
- Brand detection is a name match on the answer text, so a brand whose name is also a common word will over-count.
- Position is derived from where the mention falls in the answer, not from a ranking the model publishes.
- Sentiment is a language tally over the sentence around the mention — a three-way signal, not a nuanced read.
- Excerpts are truncated to roughly 260 characters.
- Citation sources reflect the engines that cite, so they are weighted towards Perplexity.
- The trend is built from your own stored runs, so it starts empty.
- Saved results are kept for 90 days.
Troubleshooting#
| Symptom | Likely cause | Fix |
|---|---|---|
Enter a value first. | Empty field | Type a brand domain |
Please sign in to run this tool. | Signed out | Sign in and retry |
This tool needs a paid plan… | Free plan | Upgrade — see Plans and pricing |
No measured data was found for this query. | No engine returned a usable measurement for this brand | Check the spelling, or try the domain rather than the brand word. See A tool returned no data |
| Visibility is 0 | The brand genuinely was not named in any tested answer | Read the excerpts — they show who was named instead |
| An engine row is missing entirely | That engine has no key configured on the platform | The findings name it. See AI models used across the platform |
Engines 1/5 | Only one provider is configured | See AI models used across the platform |
| The score changed with no change on your side | Each run writes its own questions and the models are non-deterministic | Read the trend across runs |
| The Claude row, or the score, moved between a run before 14 September 2026 and one after | The Claude check moved to a newer model on that date | Expected. Compare runs from after that date with each other. See the FAQ below |
| Identical numbers to two days ago | Shared cache hit — two days for this tool | Expected. See Result caching and freshness |
| The run takes a long time | It is querying several live models | Expected. See Tool limits and performance |
Hourly fair-use limit reached… | Over 100 light-tool calls this hour | Wait for the hourly reset |
FAQs#
Which engines does it actually query? ChatGPT, Gemini, Perplexity, Claude and Google's AI Overview — whichever of those has a key configured on the platform. The Engines tile shows the count out of five and names them, and the findings name any engine that did not run.
Why did my Claude result change around 14 September 2026? On 14 September 2026 the Claude check moved from Claude 3.5 Haiku, an October 2024 model, to Claude Haiku 4.5, the same generation the rest of the platform uses. The newer model knows more recent brands and words its answers differently, so the Claude row, and the visibility score built partly from it, can move between a run before that date and one after it. Nothing changed on your side. Compare runs from after the change with each other. The model behind each engine is listed in AI models used across the platform.
Does it mention my brand in the questions? No, and that is the point. The generated questions are category questions that never name you, so appearing in the answer is a genuine signal rather than a prompted one.
Where do the questions come from? If you have saved tracked prompts, those are preferred. Otherwise they are grounded in your domain's top ranked keywords and recent news headlines, then shaped into a mix of four category-recommendation, three comparison, two problem-solving and one pricing question.
How is the visibility score calculated? Appearances divided by chances, where chances is the number of questions multiplied by the number of engines queried, expressed as a percentage.
Why did my score change when nothing changed? Each run generates its own question set and the models are non-deterministic. Judge the Visibility Trend, not a single reading.
Is this the same as the free AI visibility checker? No. The free public checker scores whether a page is technically ready for AI search — crawler access, structured data, headings. This tool measures whether the assistants actually name you.
How does it relate to Brand Health? Brand Health runs this tool as one of six steps for 10 credits. Running it here alone costs 1 and gives you only the AI part.
Can I schedule it? The tool itself is run on demand. Continuous tracking with drop alerts is handled separately — see AI prompt tracking.
See also
Was this article helpful?
Thanks — feedback noted for the docs team.