How to run an AI visibility audit: a 7-step checklist
Run a real AI visibility audit: build the prompt set, sample correctly, score presence and citation rates — and know how precise your number actually is.
How to Run an AI Visibility Audit: A 7-Step Checklist
Most “AI visibility audits” are a person typing their company name into ChatGPT, reading the answer, and forming an opinion. That produces a feeling, not a measurement — and it fails the only test that matters: you can’t re-run it next month and say whether anything improved.
A real audit produces a rate. It uses a fixed set of buyer questions, samples them enough times to survive the randomness in how assistants answer, scores every response the same way, and records the conditions so the whole thing is repeatable. It also tells you honestly how precise that rate is, which is the step almost every version of this checklist skips.
Here is the procedure, the scoring, and the math on how much your resulting number can actually be trusted.
Key Takeaways
- Run each prompt five times per engine. A variance study of 12,933 brand responses found run-to-run resampling accounted for 34.8% of variance and that a single response was statistically unusable for ranking brands — with clear diminishing returns past the fifth repeat.
- Score four numbers, not one: presence rate (named), citation rate (linked), the gap between them, and share of voice against competitors.
- Repeated runs of the same prompt are correlated, so your effective sample size is closer to the number of prompts than the number of runs. At 25 prompts, expect roughly ±18pp of margin — enough to tell “invisible” from “present”, not enough to prove a 5-point monthly gain.
- Audit the source layer too. The domains AI cites for your prompts are your target list, and 26% of brands studied by Ahrefs had zero AI Overview mentions at all — absence is the normal starting condition.
- The full manual audit is ~500 responses and about 12 hours per cycle. Knowing that number up front is how you decide what to automate.
What an AI Visibility Audit Actually Measures
AI visibility is how often, and how favorably, your brand appears in answers generated by assistants. Three distinctions make the audit worth doing properly.
Visibility is not traffic. An assistant can recommend you by name with no link at all. Research from Semrush and Kevin Indig covering 3,981 domain appearances across 115 prompts and four engines found 25.1% of appearances were brand mentions with no citation — influence that no analytics tool can ever record. Auditing the answers themselves is the only way to see it.
Visibility is not your Google rankings. An Ahrefs analysis of 15,000 long-tail queries found only 12% of links cited by ChatGPT, Gemini, and Copilot appeared in Google’s top 10 for the same prompt. Your rank tracker cannot stand in for this.
Being cited is not being named. In that same Semrush study, 61.7% of appearances were “ghost citations” — the page was used as a source, but the brand was never mentioned in the answer text. Any audit that scores only one of the two will misread your position.
So the audit tracks four numbers:
| Metric | Formula | What it tells you |
|---|---|---|
| Presence rate | responses naming your brand ÷ total responses | Whether buyers hear your name |
| Citation rate | responses linking your domain ÷ total responses | Whether you’re feeding the answer |
| Ghost gap | citation rate − presence rate | Positive = you’re being used anonymously |
| Share of voice | your mentions ÷ all brand mentions in the set | Your position relative to competitors |
Step 1: Build the Prompt Set
Everything downstream inherits the quality of this list, and this is where most audits go wrong before they start.
Use unbranded buyer questions. “What do you know about [your company]” is a navigational query no prospect types. The prompts that decide revenue are the ones asked before the buyer knows you exist: “best [category] for [buyer type]”, “is [category] worth it for [situation]”, “[approach A] vs [approach B] for [use case]”, “what should [category] cost, and what’s usually padded”.
Cover the funnel deliberately. Aim for roughly a third problem-stage (“how do I fix X”), a third solution-stage (“what’s the best way to do X”), and a third vendor-stage (“who should I hire for X”). Brands are frequently present on the first and absent on the third — a pattern you can only see if you sampled both.
Size it at 15–30 for a first audit. That’s enough to establish presence or absence and small enough to actually finish. Precision at that size is limited, which we’ll quantify in step 5.
Validate before you commit. A prompt nobody asks is a wasted row. Our guide to finding out what people actually ask AI about your industry covers the five discovery methods and the demand gate to run first.
Step 2: Set Clean-Room Conditions
Conditions change answers, so fix them and write them down. Anything you don’t record, you can’t reproduce next month.
- Logged out, or in a temporary chat. Account memory and history personalize answers toward brands you’ve discussed. This is the single most common way a self-audit produces a flattering, wrong result.
- One location and language, recorded. Answers vary by market.
- Search/browsing mode explicitly on or off, and kept consistent — grounded retrieval and parametric answers behave differently.
- Same engine set, every cycle. ChatGPT, Gemini, Perplexity, and Google AI Overviews is a reasonable default; which engines matter most for you depends on your audience.
- A tight date window, ideally a few days. Models and indexes shift underneath you.
Step 3: Run Each Prompt Five Times Per Engine
This is the methodological core, and skipping it invalidates everything else.
AI answers are non-deterministic — identical prompts produce different answers on different runs. A variance-components study of 12,933 LLM brand responses, covering 20 brands across 8 languages and 3 models, decomposed exactly where that noise comes from. Within-prompt resampling alone accounted for 34.8% of variance on its stability subset, while the brand’s own underlying quality contributed 1.5%. The authors’ conclusion is blunt: ranking brands from a single query yields a generalizability coefficient of about 0.01 — statistically unusable.
Two practical rules follow. First, never trust one run. Second, stop at five — the same study found each repeat past the fifth reduced relative-error variance by roughly 0.0003, so run six onward buys you nothing.
(That study measured sentiment polarity rather than presence specifically, so treat the exact coefficients as directional. The structural finding — that run-to-run noise dominates single observations — is what carries over, and it’s consistent with what anyone running the same prompt twice sees.)
Step 4: Log Four Fields Per Response
One row per response, not per prompt. Five runs across four engines on 25 prompts is 500 rows — a spreadsheet is fine.
| Field | Values |
|---|---|
| Brand named in answer text? | yes / no |
| Your domain in cited sources? | yes / no |
| Competitors named | list them |
| Cited domains | list them |
Two optional fields earn their keep if you have the patience: position (were you first recommended, or an also-ran in a list of nine?) and framing (recommended, mentioned neutrally, or mentioned as the cheap/limited option). A brand named last with a caveat is not in the same position as one named first, and a binary yes/no hides that entirely.
Step 5: Score It — and Know How Precise It Is
Compute the four metrics from the table above across all responses. Then do the step that separates an audit from a vanity number: state your margin of error.
Here’s the trap. Five runs of the same prompt are not five independent observations — they’re five correlated samples of one question, which is precisely what the variance study demonstrates. So your effective sample size sits much closer to your prompt count than your response count. Using prompt count at 95% confidence:
| Prompts | Margin of error (rate near 50%) | Margin (rate near 20%) |
|---|---|---|
| 20 | ±21.9pp | ±17.5pp |
| 30 | ±17.9pp | ±14.3pp |
| 60 | ±12.7pp | ±10.1pp |
| 100 | ±9.8pp | ±7.8pp |
| 200 | ±6.9pp | ±5.5pp |
| 384 | ±5.0pp | ±4.0pp |
Read that honestly. A 25-prompt audit reliably distinguishes “we’re essentially absent” from “we show up about a third of the time.” It cannot support “our visibility rose from 18% to 23% this month” — that claim needs a few hundred prompts, and reporting it from thirty is how AI visibility reporting loses credibility with a CFO.
So report your first audit as a band, not a point: “presence rate roughly 10–30%, measured across 25 prompts and 4 engines in July.” Then hold the prompt set fixed so the direction of change stays meaningful even while the absolute precision is coarse.
Step 6: Audit the Source Layer
Your own scores tell you where you stand. The cited-domains column tells you what to do about it.
Aggregate every domain cited across your prompt set and rank by frequency. That ranked list is the most actionable output of the entire audit, because it is literally the set of sources the engines consult when someone asks your category’s questions. Being present on those sources is upstream of being in the answer.
Two things to look for. Which types dominate — community platforms, review sites, editorial media, or competitor-owned pages — tells you what kind of presence your category rewards. And where you’re absent from a source that appears repeatedly is a concrete, finite to-do list: a directory to get listed in, a review platform to build a profile on, a publication to earn coverage from.
This matters more than any on-site change. Ahrefs’ study of 75,000 brands against millions of AI Overview responses found unlinked web mentions correlated with AI visibility at 0.664 versus 0.218 for backlinks. Brands in the top quartile for web mentions averaged 169 AI Overview mentions against 14 for the next quartile down. The same study found 26% of brands had zero mentions — so if your audit returns near-zero, you’re in ordinary company, not a special crisis.
Step 7: Check the Two Technical Gates
Before you plan any content work off these results, confirm the engines can actually reach and read your site — otherwise you’ll be optimizing something invisible.
- Crawler access. Check
yourdomain.com/robots.txtforOAI-SearchBot,GPTBot,ChatGPT-User,ClaudeBot, andPerplexityBot, then check your CDN’s bot settings separately. BlockingOAI-SearchBotremoves you from ChatGPT’s search answers entirely, while blockingGPTBotonly opts you out of training. - JavaScript rendering. Load your key pages with JavaScript disabled. No major AI crawler executes JavaScript, so anything that isn’t in the raw HTML doesn’t exist to them.
If either gate fails, fix it first — it’s cheap and fast, and it changes what every subsequent audit measures. Our full diagnostic covers both in depth, along with the six other reasons a business doesn’t show up in ChatGPT.
How to Read Your Results
| Pattern | What it means | Where to act |
|---|---|---|
| Citation rate high, presence rate low | Ghost citations — you’re feeding answers anonymously | Brand-name proximity to your own claims and data |
| Present on problem prompts, absent on vendor prompts | Your content teaches but doesn’t qualify you as an option | Comparison, pricing, and selection-criteria content |
| Present on one engine, absent on others | Source or index concentration | Diversify the third-party sources you appear on |
| Near-zero everywhere, gates passing | No corroboration layer | Third-party mentions, reviews, listings |
| Near-zero everywhere, gates failing | Access or rendering problem | Fix step 7 before anything else |
Re-run the same prompt set monthly under the same conditions. Change the prompts and you’ve reset your baseline — keep a stable core set and add new prompts as a clearly separate group.
The Honest Limit of the Manual Audit
Twenty-five prompts, four engines, five runs each is 500 responses. At about 90 seconds to read each answer, check the sources, and log four fields, that’s roughly 12 hours per cycle — and the whole point is to repeat it monthly. The math is what pushes most teams to either shrink the prompt set until the result is too imprecise to act on, or quietly stop after the first audit.
That’s the tradeoff to make deliberately rather than by attrition: the manual audit is genuinely free and genuinely rigorous, and it costs a day and a half of a marketer’s month, every month. Automated prompt tracking exists to make the sampling continuous and the prompt count large enough that a 5-point move actually means something — which is the threshold where this stops being a curiosity and starts being a reportable metric alongside AI-referred pipeline.
Either way, run the audit once by hand first. Reading 500 answers in your own category teaches you things no dashboard will — how buyers’ questions get decomposed, which competitors AI reaches for by reflex, and how your category actually gets described when you’re not in the room.
Frequently Asked Questions
What is an AI visibility audit?
An AI visibility audit is a structured measurement of how often AI assistants name and cite your brand when buyers ask questions in your category. Unlike a one-off spot check, it uses a fixed prompt set, repeated sampling across engines, and consistent scoring — so the output is a rate you can re-run next month and compare, rather than an anecdote. A complete audit measures four things: presence rate, citation rate, share of voice against competitors, and the source layer AI pulls from for your topics.
What is AI visibility?
AI visibility is how often, and how favorably, your brand appears in answers generated by AI assistants like ChatGPT, Gemini, Perplexity, and Google AI Overviews. It is distinct from traffic: an assistant can name your brand with no link at all, so a large share of your AI visibility never appears in analytics. It is also distinct from Google rankings — only 12% of the links AI engines cite appear in Google’s top 10 for the same prompt.
How many prompts do I need for an AI visibility audit?
For a first audit, 15 to 30 unbranded buyer prompts is enough to establish whether you’re broadly present or broadly absent. Be honest about the precision that buys you: with 25 prompts, the margin of error on a visibility rate is roughly ±18 percentage points at 95% confidence, because repeated runs of the same prompt are correlated and don’t count as independent observations. That’s enough to distinguish “near zero” from “about a third” — not enough to prove a month-over-month move from 18% to 23%, which needs a few hundred prompts.
How many times should I run each prompt?
Five times per prompt, per engine. AI answers are non-deterministic, so a single response has almost no diagnostic value — a variance-components study of 12,933 LLM brand responses found within-prompt resampling accounted for 34.8% of total variance, and that ranking brands from a single query was statistically unusable. The same study found sharply diminishing returns past the fifth repeat, with each additional run cutting relative-error variance by roughly 0.0003.
Can I check my AI visibility for free?
Yes. The manual audit in this article costs nothing but time — the only requirements are accounts on the assistants you want to test and a spreadsheet. The real cost is hours: 25 prompts across four engines at five runs each is 500 responses to read and log, roughly 12 hours per audit cycle. Free automated checkers give you a fast directional read on a smaller prompt set, which is usually the right way to decide whether the full audit is worth your afternoon.
Start with the two-minute version. Run your site through the free AI visibility checker to see how you appear across ChatGPT, Gemini, Perplexity, and more — no credit card required. When you want the full prompt set sampled continuously instead of once a month by hand, start a 7-day free trial.