How to Audit Your Brand’s Visibility in AI Search, Step by Step

AI visibility audit framework showing brand citation gaps in LLM-powered search results

Author:

Ara Ohanian

Published:

October 29, 2025

Updated:

October 8, 2026

In late 2025, around 600 volunteers ran the same 12 prompts through ChatGPT, Claude and Google’s AI, nearly 3,000 runs in all. SparkToro and Gumshoe, who organized it, found that ChatGPT and Google’s AI had less than a 1 in 100 chance of giving the same list of brands twice in 100 runs, and about a 1 in 1,000 chance of giving the same list in the same order.

That changes what an audit of your brand’s AI visibility has to be. The screenshot showing ChatGPT ignoring you proves very little, and so does the one where it recommended you. Why AI search leaves brands out is covered in why AI search names your competitor and not you, and which sites ChatGPT tends to cite in the sites ChatGPT cites most. This page is the procedure: two audits, a way to score them, and a way to trace each gap to its cause.

Audit one: what AI says when someone asks about you by name

Start with the questions buyers ask once they already know your name, because errors there cost you deals that are already moving. Run prompts like these in every assistant your buyers use:

  • What is [brand], and who is it for?
  • How much does [brand] cost?
  • [Brand] vs [main competitor]: which is better for [use case]?
  • What do customers say about [brand]?
  • What are the alternatives to [brand]?

For each answer, check the facts against reality: prices, plans, features, locations, ownership, the markets you serve. Note anything outdated, anything wrong, and the tone of the description. Then look at the sources the answer cites. An inaccurate claim usually traces back to a specific page: an old pricing page you never redirected, a directory listing from years ago, a review written about a product you have since changed.

That source is the fix. Update or redirect your own outdated pages, correct your listings on directories and review platforms, and ask publishers to update stale comparisons. Correcting the record at the source is slower than correcting it in a chat window, but it is the only version that lasts.

Audit two: whether AI names you when buyers ask about the category

This is the audit most people mean, and the one most often done badly.

Build the prompt set from your buyers, not your keyword list. Take 20 to 30 questions from sales call notes, support tickets and the words prospects use in emails, then write each one several ways. People rarely phrase the same need alike: in the SparkToro study, human-written prompts for the same intent had an average semantic similarity of only 0.081. One wording per question will mislead you. If your buyers use more than one language, run the set in each. A business selling in Armenia may need Armenian, Russian and English versions, because an assistant searching in one language reads different pages than one searching in another.

Run each prompt more than once, in clean sessions, in each assistant your buyers use. The SparkToro authors suggest 60 to 100 runs per prompt for a stable average. That volume is where paid tracking tools earn their fee. A manual audit with a handful of runs per prompt gives you a direction rather than a measurement, which is still far better than one screenshot.

For every run, record five things: whether the assistant searched the web, whether you were named, whether your site was cited, which competitors were named, and which third-party pages were cited.

Score it as a share of answers, not a rank

Because the lists change on almost every run, position in the list means little. The SparkToro study makes the point with one example: in one test, City of Hope appeared in 69 of 71 ChatGPT answers, 97%, but was the top mention in only 25 of them. Its visibility was near total. Its rank looked random.

So score visibility as a share. For each prompt and each assistant, calculate the percentage of runs that named you, and the same for each competitor. Then calculate a second number: of the answers that named you, how many cited your own site. The first number tells you whether you are in the consideration set. The second tells you whether your own pages are doing the work or someone else’s are. Compare both with your competitors rather than with an abstract benchmark, since no reliable industry benchmark exists yet.

Be skeptical of any tool that reports an AI “ranking.” The study’s authors put it bluntly about AI recommendation lists: “Thinking of them as sources of truth or consistency is provably nonsensical.”

Trace each gap to its cause

A low score is a symptom. The cause decides the fix, and your run log already tells you which one you have.

If the assistant searched the web but your pages never appeared among its sources, check access first. OpenAI’s crawler documentation separates OAI-SearchBot, which powers ChatGPT search, from GPTBot, which gathers training data, and a site can allow one while blocking the other. Perplexity runs PerplexityBot for its search results, according to its documentation. Check robots.txt, your firewall rules and whether the content you want quoted is in the page’s HTML rather than loaded by JavaScript. Then check whether you have a page that answers the exact narrow question at all. Fixing access and indexation is the groundwork of search engine optimization built for AI answers as well as Google.

If the cited sources are third-party lists and reviews that leave you out, those pages are your outreach list. Sort them by how often they appear across runs and work from the top.

If the assistant answered without searching and still left you out, the gap is in what the model learned during training. That changes only as the web’s coverage of you grows and new models are trained, so treat it as a long-term result of the other fixes.

And if you were named but described wrongly, go back to audit one and fix the source.

Finally, check that the effort reaches your business. In your analytics, segment referrals from chatgpt.com, perplexity.ai, gemini.google.com and copilot.microsoft.com, and in your server logs, watch for visits from the search crawlers. Mapping all of this against competitors, and ranking the gaps by what they cost you, is what an analysis of where your brand shows up online, and where it doesn’t covers.

Rerun the same prompt set on a fixed schedule, monthly if AI answers matter to your pipeline, and change the set only on purpose. AI visibility is a share you defend over time, not a position you win once.

Frequently Asked Questions

How do I check my brand’s visibility in AI search?

Run two sets of prompts across the assistants your buyers use. The first asks about your brand by name, to check whether the AI describes you accurately. The second asks the category questions buyers ask before they choose, to check whether you are named at all. Run each prompt several times in clean sessions, record whether you were named and cited and which competitors appeared, and score the results as the share of answers that named you.

Why does ChatGPT mention my brand sometimes and not other times?

AI assistants generate a fresh answer each time, and when they search the web, small differences in the searches they run change the pages they read. A SparkToro and Gumshoe study found ChatGPT and Google’s AI had less than a 1 in 100 chance of returning the same list of brands twice. That is why a single test proves little. Measure how often you appear across many runs rather than whether you appeared once.

How often should I audit my brand’s AI visibility?

Monthly if AI answers influence your sales pipeline, and at least quarterly otherwise. Use the same prompt set each time so changes in your results reflect changes in the AI answers, not in your questions. Run the brand-accuracy prompts after any change to pricing, products or positioning, since outdated pages and listings are where wrong descriptions tend to start.

Are AI visibility tracking tools worth paying for?

They can be, mainly because they run prompts at a volume that manual audits cannot match, and stable averages need many runs per prompt. Judge them on method. Ask how many times each prompt runs, how prompts were written and whether results are reported as a share of answers. Be wary of any tool that reports a fixed ranking position, since AI answers rarely repeat the same order.