
An AI discoverability self-audit tests technical retrievability, answer appearance, citation, and accurate representation separately using fixed prompts, documented evidence, entity checks, and cautious retesting.
An AI discoverability self-audit sounds as if it should answer one question: can AI find you?
But that question hides several different questions. Can an AI system retrieve your content? Does your subject appear in its answer? Does the answer cite a source? Is what it says accurate? A page can be reachable but never appear. A subject can appear without a citation. A cited answer can still be wrong.
So the useful audit is not a single visibility check. It is a set of separate observations, gathered under known conditions and compared over time. The result will not tell you how every AI system sees you; it will show what happened for the prompts, systems, places, languages, dates, and other conditions you tested.
Start by defining four outcomes:
These outcomes are related, but they are not interchangeable. Machine-mediated search can retrieve a page without citing it, summarize it without a visible source, or use several sources in a recommendation. Citation is therefore one observable result, not a synonym for AI discoverability.
This distinction keeps the audit from becoming misleading. If the site is reachable but the subject never appears, you have a different problem from one in which the subject appears with the wrong pricing. Likewise, a recommendation without a link may be useful even though it fails an AI citation audit; a citation attached to a false claim is not a win.

Next, write down the scope. Name the subject being audited, the geography and language that matter, the user needs you intend to test, and the AI-mediated answer environments likely to be used. Do not begin with a fixed list of platforms and assume it applies everywhere. Choose systems according to likely user behavior, market, and use case. Possible environments include ChatGPT, Gemini, Claude, Perplexity, Microsoft Copilot, and Google AI Overviews or AI Mode, but platform availability and behavior change.
Then create a canonical fact set: the current facts against which answers will be checked. Record the canonical name, official domain, description, category, audience, products or services, locations, pricing model, important features and limits, relationships to people or parent organizations, and any other claims that would materially affect a user’s decision. For each fact, preserve its official source and collection date.
This fact set matters more than it first appears. Without it, “accuracy” becomes an impression. With it, you can mark a claim as correct, incomplete, outdated, unsupported, or false.
Finally, state what the audit cannot establish. A manual prompt sample is a local benchmark, not a statistically representative view of all queries or users. It cannot produce a universal AI search visibility score, prove that a visible source caused an answer, or guarantee future inclusion, citation, or accuracy. Its job is narrower and more useful: reveal observed weaknesses, recurring patterns, and priorities worth investigating.
The prompt set is the center of the audit. If it changes casually from one run to the next, the comparison means little.
Build prompts from language real users or buyers employ, not from internal marketing terms. Organize them into clusters that test different kinds of discovery:
A direct prompt might ask what the subject is. A category prompt might ask for suitable options without naming it. A comparison prompt might test whether the system distinguishes it from a competitor. The point is not to hit a prescribed prompt count. It is to cover distinct ways the subject should reasonably appear—and distinct ways the system could fail.
Lock a core set for the test cycle. Preserve the wording exactly, including constraints and requested geography. If you want to explore variants, label them as variants rather than quietly replacing the baseline prompt.
Before running anything, decide which conditions will remain fixed and which will vary. At minimum, consider:
Not every environment lets you control or even observe every variable. Record that uncertainty instead of pretending it is absent. A logged-in answer may be personalized; a continuing conversation may depend on earlier turns; location can affect retrieval. These are limits on interpretation.
Run the same core prompts across the relevant environments. Repeat the important prompts, because answers can vary by run, wording, time, engine, and location. Repetition does not reveal the complete behavior of a model, but it keeps you from mistaking one unusual answer for a stable pattern.
Use a structured evidence log. Each row should preserve:
Keep the raw answer. Do not save only your summary. As practical audit guidance notes, capturing the exact generated text is necessary for later analysis. A summary may erase the very wording that made a claim misleading.
Also preserve URLs before grouping them by domain. A homepage, pricing page, help document, review page, and comparison article on the same domain can support different claims and call for different fixes. Domain-level patterns are useful later; the page-level evidence comes first.
You now have something more modest than a universal measurement, but much better than casual testing: a reproducible record of what selected systems did with selected prompts under documented conditions.
Before interpreting answer-level results, check whether the source material was available to be retrieved.
Choose representative production URLs, not merely a test page that happens to work. Include priority pages and each important template: the home or entity page, product or service pages, pricing, locations, biographies, documentation, and any pages containing the canonical facts used in the audit.
For each representative URL, inspect:
noindex signals;Run these checks against production responses and record the requesting user agent, status, redirects, canonical URL, rendered text, and important fields successfully extracted. A normal browser visit is not enough; a CDN or bot-management layer may serve a crawler a different response.
Crawler names also require care. Providers may separate training crawlers, search crawlers, and user-triggered fetchers. OpenAI, for example, documents independent controls for OAI-SearchBot and GPTBot. Allowing one does not prove access for another provider, crawler, or use case. General search crawlability is useful evidence, but it is not direct proof of access by a particular answer system.
Robots rules are not a complete exclusion test either. Google’s robots.txt documentation explains that robots.txt controls crawling rather than guaranteeing that a URL will disappear from search; a blocked URL may still be discovered through external links. Crawlers may also interpret or obey rules differently. Record the exact rule, user agent, and tested result rather than writing “AI bots allowed” as if it were universal.
Check raw HTML or view source on the important pages. Confirm that titles, headings, navigation, internal links, metadata, descriptions, pricing, product or service details, and other critical facts remain available without client-side JavaScript. Some automated systems can render JavaScript and some cannot do so consistently, so the safe audit rule is not to assume either. It is to verify what exists before rendering and what appears afterward.
Then inspect server logs. Logs can show that a named requester reached a URL, when it arrived, and which status code it received. They can reveal sudden drops, repeated errors, or changes after a deployment. But the absence of a request is inconclusive: it does not by itself tell you whether the cause was crawler policy, lack of discovery, platform behavior, or something else.
Record the result as a technical access finding. Passing means the page was available under the tested conditions. It does not mean that the page was indexed, selected for a prompt, considered authoritative, cited, or represented correctly. Crawlability, answer visibility, and citation accuracy should be reported separately.
That is the first reveal in the audit: the technical check does not tell you whether you are discoverable. It tells you whether one possible obstacle can be ruled in or out.
The next question is whether the subject is described in a way that can be matched consistently across pages and records.
Begin with a structured-data inventory on representative URLs and templates. Record the types, properties, identifiers, and referenced URLs actually present, including markup added after JavaScript rendering. Look for persistent identifiers such as @id, and inspect the stated relationships among organizations, people, products, pages, and outside references. Connect records only when they refer to the same real-world thing.
Validation has two separate parts. First, confirm that the markup is actually present in the production page or rendered output. Second, validate it. Schema.org syntax validation and search-feature eligibility validation answer different questions; a record may be valid Schema.org markup but ineligible for a particular search feature. Record errors separately from warnings.
Then compare the markup with the visible page. Look for stale prices, hidden claims, duplicated records, conflicting names, wrong entity references, invalid syntax, and properties that disagree with what a user sees. Valid code containing false or old facts is still bad data.
Now widen the check beyond markup. Inventory controlled site surfaces and relevant independently controlled references. Compare:
Classify who controls each source. An official profile is not independent corroboration merely because it sits on another domain. Nor should you create a profile or identifier solely to fill an audit checklist. No external platform is universally required or known to be used by every answer system.
Flag contradictions, stale records, ambiguous matches, and mismatches as data-quality findings. A changed official URL still referenced elsewhere is a finding. So is a product code that differs across pages, or a founder relationship stated two ways. These inconsistencies may justify cleanup, but they are not automatically “citation losses” or ranking failures.
This is where entity optimization is easiest to overstate. Consistent records and valid structured data make identity signals clearer and easier to audit; they do not prove that an AI system retrieved the page or used the markup. Schema can provide clearer signals without guaranteeing visibility.
Keep answer testing separate. When you test whether an AI system identifies the entity correctly, preserve the prompt, system, date, answer, cited URLs, and entity match. An association between cleaner records and a later answer change may be worth studying, but the before-and-after observation does not reveal the mechanism.
Now return to the captured answers and inspect what users actually saw.
For every answer, distinguish a mention from a citation. A mention means the subject appears in the answer. A citation means the system exposes a linked or attributed source. Record both, along with the subject’s role: was it merely named, described, compared, shortlisted, recommended, or ruled out?
Preserve every exact cited URL before aggregating recurring domains. Then map each material answer claim to the visible cited page where possible. Open the page and ask a plain question: does this source actually support this claim?
Citation presence is not proof. A linked page may be stale, inaccessible, irrelevant, or too weak to support the wording attached to it. For each visible or likely source page, assess:
Next, compare every material claim with the canonical fact set. Check identification, category, intended audience, products or services, pricing, locations, features, capabilities, limitations, relationships, and other current details. Mark not only outright errors but partial truths that create a false impression.
Look for omission too. An answer may contain no false statement yet still leave out a limitation that changes who the product fits, or fail to answer a subject-specific question despite the information being available. There is no universal threshold for a material omission, so explain why the missing fact matters.
Test ambiguity directly. Does the answer conflate the subject with a similarly named organization, a competitor, a founder, another location, or an unrelated product? Does it place the subject in the wrong category or use case? Repeated misclassification is evidence of a representation problem; it is not, by itself, proof of which page or signal caused it.
Also separate factual correctness from framing. An answer may contain individually true claims while giving harmful or misleading emphasis. Record negative or biased framing when it changes the practical meaning, but do not confuse negative sentiment with inaccuracy. A critical answer can be accurate; a flattering answer can be false.
When no citation appears, do not assume the answer had no source. Extract repeated or material claims and compare them with plausible source pages. Label every such source inferred, not proven. The same rule applies when a follow-up asks the system for sources: preserve what it returns, but do not treat the reply as definitive evidence of the internal path that produced the original answer.
If a citation cannot be opened, verified, or matched to the claim, record that limitation. If the source contradicts the answer, separate the claim from the citation and mark it unsupported or requiring verification.

Finally, compare repeated runs and environments. Track changes in mention rate, citation rate, prominence, recurring sources, source mix, wording, competitors, and claim accuracy. One answer can expose a serious error. It cannot establish that the error—or the success—is stable.
A useful report preserves the parts instead of hiding them in one number.
Report findings by prompt cluster and platform across the main components: technical access, entity clarity, answer appearance, citation quality, and representation accuracy. Include the raw evidence and your confidence in each finding. A repeated error supported by the answer, canonical facts, and source page deserves more confidence than a one-off omission under poorly controlled conditions.
You may use a coarse local rubric—such as absent, weak, decent, or strong—to help sort work. If you do, define each label, retain the component observations, and call it an internal decision aid. Sources propose incompatible scoring models, while others acknowledge that no general benchmark for a good AI visibility score has been established. A precise composite can make uncertain evidence look settled.
Prioritize each finding using four considerations:
This tends to put a repeated pricing error, a production access block, or systematic entity conflation ahead of a one-off low-prominence mention. Assign an owner and due date to every accepted action. Group the backlog into quick fixes, structural work, and ongoing measurement so that easy corrections do not obscure larger template or data problems.
Retesting should be designed when the finding is logged, not after the fix ships. Preserve the original baseline: prompt wording, platform and mode, market, language, user state where known, date, raw response, cited URLs, competitors, accuracy assessment, and run-to-run stability. After implementation, rerun the same approved prompts under conditions as comparable as you can make them.
Recheck the affected URLs too. Changes to templates, headings, structured data, robots rules, pricing, brand names, migrations, navigation, authentication, or CDN behavior can alter technical and semantic conditions beyond the page you intended to edit.
Do not claim causation from a better retest. The change may have helped; model updates, source changes, competitor activity, location, personalization, or ordinary response variation may also have contributed. Report the observed difference and the plausible explanations.
There is no universal cadence. Use lighter recurring checks where answers or facts change quickly, deeper periodic audits for the whole prompt and page set, and event-triggered retests after material technical, content, brand, or migration changes. Keep the fixed core prompts even when you add exploratory ones.
The real output of an AI discoverability self-audit is not a score. It is a chain of evidence: what was reachable, what appeared, what was cited, what was said, and under which conditions. Once you have that, the next question becomes much sharper—not “Are we visible to AI?” but “Which observed condition should we change, and what would count as credible evidence that it improved?”
If you want to set up a deep AI audit that runs regularly and gives you rich insights into what AI is saying about you, there are tools out there to make the process a little more painless. Check out Onsomble which allows you to set up an audit in minutes.
Most robots.txt guides for AI crawlers are written for publishers who want to block the bots. This is the opposite: how to check if your site is accidentally invisible to ChatGPT, Claude and Perplexity — and the exact lines to paste to fix it.
Most synthesis guides are built for academics writing literature reviews. Here's a workflow for knowledge workers who need to turn 20+ sources into a deliverable with a deadline.
RAG and fine-tuning solve different problems. Here's how to decide which one your project actually needs.
AI patterns, workflow tips, and lessons from the field. No spam, just signal.