How to Track Brand Mentions in AI Answers Every Month
Key Takeaways
- Tracking means running the same prompts, on the same engines, on the same day each month. A changed prompt list produces a changed result for reasons that have nothing to do with your marketing.
- Twenty prompts is enough: eight category, six comparison, four problem and two direct brand prompts. Freeze the list for at least two quarters before revising it.
- Record five columns per answer: mentioned, cited, position in the answer, how you were described, and which competitors were named.
- AI answers vary between runs, so a single mention appearing or disappearing is noise. Only a change that holds across two consecutive months, or across repeated runs of the same prompt, is a signal.
- Markty AI runs four scans a month across ChatGPT, Gemini and Perplexity, keeps the history and benchmarks competitors, which removes the discipline problem that kills most manual tracking.
Tracking AI mentions means running the same prompts, on the same engines, on the same day each month, and recording the same five things every time. The one-off audit tells you where you stand. Only a repeated panel tells you whether anything you did worked, and without that, every content decision you make about AI visibility is a guess.
Build the panel once
Twenty prompts, written as a real customer would type them, split roughly like this:
Eight category prompts. "Best [category] for [your customer type]", "affordable [category] for a small team", "[category] that does [specific job]".
Six comparison prompts. "[Competitor] alternatives", "[competitor A] vs [competitor B]", "cheaper alternative to [leader]".
Four problem prompts. The job the customer is trying to do, with no product language at all.
Two direct prompts. "What is [your brand]", "is [your brand] any good".
Then freeze the list. Not for a month, for at least two quarters. The single most common way this exercise fails is a well-meaning revision of the prompt set, which resets the baseline and quietly destroys the comparison you were building.
The two direct prompts are worth keeping even though they always mention you. They are how you detect a description drifting, and a drifting description is a problem you want to catch early rather than after it has been repeated everywhere.
The five columns
For every answer, record:
Mentioned. Was your brand named in the answer text? Yes or no.
Cited. Was your site used or linked as a source? Yes or no. These fail independently and must never be merged into one column.
Position. First named, in the middle, or last. Being one of three is not the same as being one of eight.
Description. The words used about you, copied verbatim. This is the column that catches positioning problems nothing else will show you.
Competitors named. Who else appeared. Over six months this becomes a share-of-voice picture nobody else in your market has.
Sixty rows a month across three engines. About an hour, and it fits in a spreadsheet.
Telling signal from noise
These systems are probabilistic. Ask the same question twice and the three brands named may not be the same three. If you treat every fluctuation as a result you will spend the quarter chasing ghosts.
Two rules keep this honest:
Run the important prompts three times. Record how many of the three named you: zero, one, two or three. That converts an unstable yes-or-no into a consistency measure, which is a real number that moves in a real direction.
Require two consecutive months. A change that appears in one month and vanishes the next was variance. A change that holds is a signal. This rule alone will save you from most of the wrong conclusions available in this discipline.
Reading the result
Three summary figures are worth carrying forward each month.
Mention rate. Answers naming you, divided by total answers. The headline number.
Mentions-to-citations ratio. If you are cited far more often than you are named, your content is working and your identity is not, which is a specific and fixable problem rather than a general one.
Per-engine split. Never average the engines together. A strong Perplexity result and a zero on Google's surfaces is a completely different diagnosis from an even spread, and the two require opposite work: one is a content problem, the other is an entity and reviews problem.
When to act on it
Give any change 30 to 60 days before judging it. Pages need recrawling and answers refresh on their own schedule, so a fix made in week one may not appear until the second month's panel.
Act when a decline holds for two months, when a competitor appears consistently in prompts where you never do, or when the description column starts saying something you did not choose. The last one is the most valuable early warning in the whole exercise and the one almost nobody tracks.
Or have it run for you
Markty AI scans ChatGPT, Gemini and Perplexity four times a month, keeps the history so month-over-month comparison is automatic, benchmarks competitors on the same prompts, and converts the gaps into the content most likely to close them. The manual panel above is the same method; the difference is that a scheduled scan still happens in a quarter when everything is on fire.
Whichever way you run it, the discipline is the same: identical prompts, separate columns for mentions and citations, monthly, and no revisions to the panel until you have two quarters of history worth comparing.