An AI visibility audit measures how AI assistants — ChatGPT, Perplexity, Gemini, Google AI Overviews — talk about your store when shoppers ask them for recommendations. It answers three questions: do the assistants mention you, do they cite your site as a source, and what do they say instead when they skip you. In 2026 audits range from free automated scans to white-glove engagements at $1,200–$7,500.
We run these audits for a living, and this year we ran one on ourselves. The results were uncomfortable enough to publish, so this guide includes our own numbers as a worked example.
What an AI visibility audit actually checks
An AI visibility audit asks AI assistants the questions your buyers ask, then records what comes back. A serious audit separates three query classes: brand queries ("what is [your store]"), buyer queries ("who sells lab-grown emerald rings with GIA certificates"), and comparison queries ("best supplement brands for postpartum recovery"). For each one it logs two numbers per platform: mention rate (did the assistant name you) and citation rate (did it link your domain as a source). It also records which competitors filled the space you didn't. In August 2026 we ran this exact protocol on our own company: 10 prompts, 8 surfaces, 290 live calls, $4.49 in API spend. Brand queries scored a mention rate of 1.0 on all eight surfaces. Every buyer query scored 0.0. That gap is the single most common pattern we see, and no single-session spot check will reveal it.
The surfaces matter more than most people expect. ChatGPT answering from memory, ChatGPT with web search, the ChatGPT shopping app, Perplexity's fast mode, Perplexity's research agent, Google AI Overviews, Google AI Mode and Gemini are eight different systems with different retrieval behavior. In our own August measurement, Perplexity's agent cited sources on 50 of 50 calls while ChatGPT's memory surface cited zero sources on 50 of 50 — same company, same prompts. An audit that checks "ChatGPT" and stops has measured one room of the building.
Why one chat session is not an audit
Asking ChatGPT about your brand once and screenshotting the answer is the audit equivalent of weighing yourself with one foot on the floor. Assistants sample their answers: the same prompt can return different brands on different runs. A defensible audit repeats each prompt and reports the rate, not the anecdote, plus a stability score that says how reproducible the result is. Our own baseline run in August 2026 repeated 80 cases and landed at 0.985 stability, which is what let us say "buyer queries are zero" as a fact rather than a bad day. The other thing repetition buys you is a frozen panel: lock the exact prompt set, rerun it monthly, and the delta between runs measures your progress instead of measuring changes in your own questioning. Without a frozen panel, month-two numbers are not comparable to month-one numbers and any claimed improvement is unverifiable.
What it costs in 2026
Prices cluster in three bands, and the bands map to how much human judgment is involved. Free automated scans exist as lead magnets: Alhena AI, for example, has offered a free AI visibility audit with a 48-hour turnaround. Self-serve monitoring platforms — Peec AI, Otterly.AI and similar — sell monthly subscriptions that track mentions and citations on a schedule; across the market these run from roughly $99 to several hundred dollars per month depending on prompt volume and seats. Human-reviewed audit engagements, where a person adjudicates every AI answer instead of trusting keyword matching, start around $1,200 one-time (that is our beta price) and reach $4,000–$7,500 for enterprise scopes with competitor benchmarking. The honest disclosure: automated mention-counting misclassifies answers often enough that we built a mandatory human adjudication step into ours. If a report ships with zero human review, price it accordingly.
Who offers one
Five names come up repeatedly in 2026, ours included, and they are genuinely different products. Peec AI sells AI search analytics for marketing teams: visibility, position and sentiment tracking across ChatGPT, Perplexity and Gemini, with competitor benchmarking. Otterly.AI sells AI search monitoring — brand mentions, website citations, prompt research — positioned for SEO teams. Profound sells enterprise answer-engine optimization with monitoring plus AI-generated content workflows. Alhena AI leads with a free audit and sells AI shopping assistants and support agents for e-commerce. Arbling — us — is merchant-side: we audit AI visibility the same way (that is the report you are reading about), but the product under it repairs the product data itself and republishes it to the surfaces agents read, then tracks AI-attributed revenue. Monitoring tells you where you stand. The repair loop is the part where the number moves.
The selection question is simple: if you need a dashboard for a marketing team, the monitoring tools fit. If you are a merchant and the audit finds broken product data, you want the audit wired to the thing that fixes it.
How to read an audit report
A useful report separates what was measured from what it means, and shows its raw evidence. Ask for four things before you pay. First, the per-surface breakdown: an aggregate "AI visibility score of 62" hides the fact that Perplexity loves you and ChatGPT has never heard of you, which are two different problems with two different fixes. Second, the evidence file: every claim in our reports traces to a logged API response with a timestamp, and any vendor who measured something real can show the same. Third, the competitor fill: when the assistant skipped you, who did it name instead? That list is your actual competitive set in AI search, and in our experience it rarely matches the competitor list the merchant walks in with. Fourth, the fix list ranked by effort: schema and data fixes a developer ships in a day should be separated from content work that takes a quarter.
What happens after the audit
The audit is the diagnosis. The value is in the loop that follows: fix the highest-impact gaps, wait for recrawl, remeasure on the same frozen panel, and compare. On the fix side, the work is usually unglamorous — product data completeness, server-rendered content that AI crawlers can actually read, consistent entity facts across every page you own (we found our own site publishing three different customer counts simultaneously; the audit caught it). On the measurement side, monthly is the practical cadence: AI surfaces re-index slowly enough that weekly deltas are mostly noise. Adobe's retail data explains why merchants bother: AI-referred retail traffic is up over 1,300% since 2024, and those visitors convert 54% better than average as of May 2026. The channel is real. The only question an audit answers is whether it can see you.
If you want the version of this we run for clients — 8 surfaces, repeat sampling, human adjudication, evidence files included — start at arbling.com. If you want to try a rough version yourself first, our guide on tracking whether ChatGPT recommends your store walks through the manual protocol.