AI Visibility Score: How It Is Calculated and What a Good Score Looks Like
Key Takeaways
- An AI visibility score estimates how often, and how favourably, AI assistants name your brand when answering questions in your category. It is a diagnostic instrument, not a leaderboard position.
- Any credible score is built from four separate dimensions: brand visibility, per-model visibility, keyword coverage and in-answer context. Collapsing them into one number is convenient, but the diagnosis lives in the four.
- You can calculate a rough score by hand in an hour: twenty prompts, three engines, one point per mention, half a point for a citation without a mention, expressed as a percentage of the maximum.
- There is no universal pass mark. The two readings that matter are your score against direct competitors on the same prompts, and your own score last month.
- Markty AI computes this continuously as an AI Score out of 100, scanning ChatGPT, Gemini and Perplexity four times a month and benchmarking the result against competitors.
An AI visibility score is a single number that estimates how often, and how favourably, AI assistants name your brand when they answer questions in your category. It exists for one reason: AI visibility is otherwise an anecdote. Somebody asks ChatGPT something, sees a competitor's name, and the conversation about it goes nowhere because there is no number to argue with.
What the score is actually made of
Every credible implementation is built from the same two ingredients.
Topic coverage. Out of all the topics that make up your category, how many produce an answer that includes your brand? A business named in answers about scheduling but never in answers about analytics, reporting or pricing has narrow coverage, and narrow coverage caps the score no matter how strong the mentions are where they exist.
Mention consistency. When you are named, does it hold? Run the same prompt three times and you will not always get the same three brands. A name that appears in one run out of three is a weaker signal than a name that appears in all three, and a serious score treats those differently.
Add the engines and you have the grid: coverage across topics, consistency across runs, measured separately per model.
The four dimensions worth keeping separate
One number is useful for a meeting. The diagnosis lives in the four underneath it.
Brand visibility. Are you named at all, anywhere in the category? This is the coarse pass or fail.
Per-model visibility. How you score on ChatGPT, Gemini and Perplexity individually. These genuinely differ, and a single blended figure hides the most actionable finding you will get, because a model that reads the live web and a model that leans on brand familiarity need opposite fixes.
Keyword coverage. Which queries and phrases you appear alongside, and which ones you never touch. This converts the score into a content plan.
In-answer context. What the answer says about you when it names you. Being called the affordable option, the enterprise option or the one for beginners is a positioning outcome, and it is the dimension most tools ignore entirely.
In-answer context is the one to watch. A rising score with a drifting description means models are learning about you from sources you did not write, and that is worth knowing before it hardens.
Calculating a rough score yourself
An hour of work gives you a defensible baseline.
Write twenty prompts a real customer would ask. Roughly half category prompts (best X for Y), a quarter comparison prompts (alternatives to the market leader), a quarter problem prompts (how do I solve Z).
Run each in a fresh session in ChatGPT, Gemini and Perplexity, with memory or personalisation off. Sixty answers.
Score each answer: one point if your brand is named, half a point if your site is cited but your name is absent, zero otherwise.
Divide by sixty and express as a percentage. Record the three per-engine subtotals separately, because that is where the instruction is.
The half point matters. It separates "the model has never heard of you" from "the model reads you and credits someone else", and those two problems have nothing in common. The second is covered in the twenty-minute audit, which is the lighter version of this exercise.
What a good score looks like
There is no pass mark, and any tool that implies one is selling the number rather than the diagnosis. Scores are normalised against a category and a prompt set, so a 40 in a crowded market with three entrenched incumbents can represent a better commercial position than a 70 in a niche with two competitors and no demand.
Two readings are worth having:
Against competitors, on the same prompts. The only fair comparison. If you and three rivals are measured on one identical prompt set, the ranking is real.
Against yourself, last month. Direction beats altitude. A score moving from 22 to 31 over a quarter is a working programme. A flat 55 is a plateau, and usually an off-site one.
One more figure is worth tracking alongside the score, because it is more diagnostic than the score itself: the ratio of mentions to citations. If AI engines cite your site twenty times and name you three, your content is strong and your identity is invisible. That ratio moves before the score does, which makes it an early signal that the work is landing.
What the score does not tell you
It does not tell you about volume. Nobody knows how many people asked an assistant about your category last month, and any tool claiming to is estimating. It does not tell you about conversion either, since an AI mention is a recommendation, not a click, and it often shows up in your analytics as direct traffic weeks later.
Treat it as a diagnostic instrument, in the same family as a share-of-voice figure. Useful for direction, misleading if you optimise for the number itself.
Having it measured for you
Markty AI computes this as an AI Score out of 100, scanning ChatGPT, Gemini and Perplexity four times a month, breaking the result into brand visibility, per-model visibility, keyword coverage and in-answer context, and benchmarking the whole picture against competitors. It also converts the keyword coverage gaps into the content most likely to close them, which is where the score stops being a report and starts being a plan.
The number is not the point. Knowing which of the four dimensions is dragging is the point, because each one is fixed by different work: identity sentences on your own pages, presence off-site, comparative content, or consistency across every profile you own.