On this page
An AI visibility audit tells you whether ChatGPT, Google AI Mode, Perplexity, Gemini, and other answer systems mention, cite, or recommend your business when a customer asks a question you should win.
Give me an afternoon, a spreadsheet, and a private browser window. By the end, you will know how often AI names you, which competitors it names instead, the gaps you have and which sources it trusts, and whether your result survives a second or third run.
When we launched the free AI Visibility Scorecard, 1,500 businesses ran it in the first week. Eighty-nine percent were invisible to AI. Most had no idea.
You cannot fix a visibility problem you have never measured. Start here.
What is an AI visibility audit?
An AI visibility audit is a practical assessment of how AI search systems represent your business. Use the first pass as an evaluation of your baseline. It answers one question: when a potential customer asks AI platforms who to buy from, hire, or use, does your name come up?
It is the AI-search equivalent of checking Google rankings, with two differences. There is no fixed position. An AI answer is a paragraph, and you are either in it or you are not. The answer can also change when you run the same prompt again, so you measure a rate of appearance rather than a permanent spot.
The audit has four jobs:
Choose the prompts your customers use.
Run those prompts across the models that matter to your market.
Log the answer, the brands named, and the sources cited.
Diagnose the gap and rerun the same test later.
Keep three outcomes separate. A mention, a citation, and a recommendation are different signals.
All help you understand the sentiment, authority, and presence your brand has when AI crawlers assess your content and online findings.
Mentions, citations, and recommendations measure different things
A mention is your name appearing anywhere in the answer. “Other real estate agents include Brand A and Brand B” puts you in the paragraph. That is all it proves.
A citation is a link to a page attached to the answer as a source. The model found a page it could use and chose to show that page to the reader. The citation may point to your website, your Zillow profile, Realtor.com, a local publication, or another page that describes you.
A recommendation is the commercial outcome. The model names you as a pick, gives a reason, and connects you to the customer’s situation. Weight recommendations more heavily in your notes because a passing mention and a qualified referral are not equivalent.
A source can influence an answer without getting a visible link
A source is any page the system retrieves or reads while building an answer. A citation is the subset of those pages it shows as a link.
Ahrefs analyzed 1.4 million ChatGPT 5.2 prompts from February 2025 and reported that ChatGPT cited about half of the URLs in its retrieval pipeline. Its study also found that 88.46% of the URLs in the search ref_type were cited, while dedicated Reddit, YouTube, and academia channels behaved differently. The data is observational, not a published OpenAI ranking formula, but the operational lesson is clear: being retrieved and being cited are separate events. Read the Ahrefs study.
If ChatGPT recommends you and links to your G2 profile, G2 received the visible citation while you received the recommendation. Log both. The pair tells you where the system found evidence about your business and which page it trusted enough to expose.
How to find your current AI visibility score
Run five prompts first. Ten minutes are enough to tell you whether the full audit will measure a small gap or a crater.
Open ChatGPT in a private browser window, log out, and ask questions that match your market and the topics you want to be in the responses for.
For a Newton real estate agent, that might be:
“What is Newton Harbor Realty?”
“Who is the best listing agent in Newton, Massachusetts for a seller with a $2 million home?”
“I am buying my first home in Newton with a $500,000 budget. Which agent should I speak to?”
“Is Newton Harbor Realty a good choice for selling a home?”
“Newton Harbor Realty reviews.”
To effectively run an AI visibility audit, you should start by analyzing the prompts users commonly enter, as these reveal the questions and concerns your AI needs to address most accurately. Replace the example business and market with your own. Repeat the five prompts in Google AI Mode and Perplexity, then save the answers rather than trusting your memory.
A wrong description, a discontinued service, or “I do not have information about that” points to an entity problem. AI does not have a stable picture of who you are.
A correct branded answer followed by competitors on category and problem prompts points to a corroboration problem. AI knows you exist but has less evidence for recommending you.
Showing up across all fifteen results is useful, but it is still a baseline. Run the full audit before you call the result durable.
Decide What Goes Into the Audit
An audit is only as useful as its prompts. Define the customer question, the models, the location, and the exact prompt set before you open the spreadsheet.
Choose a customer question narrow enough to score
“Real estate agent” is not a target. “Luxury real estate agent for homes above $10 million in Austin” is a target. The tighter definition gives the model a clearer connection between your business, your customer, your service, and your market.
Write down three to five target questions. Each should combine what you do, who you do it for, and where the service applies. Use the words buyers and sellers use, not the language on your internal strategy deck.
If your team calls the service a “revenue intelligence platform” while buyers ask for “sales forecasting software,” the customer language belongs in the prompt list. Your existing keyword research is a useful starting point. If you have not done it, start with the difference between AI SEO and traditional SEO before you build a tracking system.
Test the models your customers can reach
Start with four surfaces. Add Claude, Copilot, and Grok after the process is stable, or move one of them earlier when your customers use that product heavily.
| Model or surface | Why include it | First-pass setup |
|---|---|---|
| ChatGPT | Tests a conversational AI-search experience with web sources. | Private window, logged out. |
| Google AI Mode and AI Overviews | Tests the generative layer attached to Google Search. | Record market and device context. |
| Perplexity | Shows source links inline, which makes source mapping easier. | Use the same location and prompt wording. |
| Gemini | Adds a Google ecosystem surface to the comparison. | Keep the account state consistent. |
Google says AI Overviews and AI Mode may use query fan-out, where several related searches across subtopics and data sources help form one response. Google also says the same foundational SEO practices apply: pages need to be indexed and eligible for a snippet, and there are no special AI-specific technical requirements. Read Google’s current guidance for AI features.
Build the prompt list around four customer situations
Your prompt list is the audit. Build 20 to 30 prompts, with more category and problem prompts than branded prompts. Those are the prompts closest to a new customer choosing who to call.
| Prompt category | Example | What it tests |
|---|---|---|
| Branded | “What is Newton Harbor Realty?” | Entity recognition and factual accuracy. |
| Category | “Who are the best listing agents in Newton for 2026?” | Recommendation for the service and location. |
| Comparison | “Compass versus Keller Williams agents in Newton.” | Whether you enter the consideration set. |
| Problem | “I need to sell my Newton home within 60 days. Who should I call?” | Match between your footprint and a real situation. |
Add branded prompts such as pricing, reviews, service area, and years in business. Add category prompts for buyer’s agents, listing agents, probate sales, relocation, and luxury homes. Add comparison prompts that name the brokerages your local clients already consider. Add problem prompts with a time limit, neighborhood, property type, budget, or other constraint.
Write the prompts the way a real person types them. “Best listing agent Newton MA” and “I am selling a three-bedroom house in Newton and need someone who knows the local market” are both valid. One is compressed. The other is conversational. Your audit should see both.
Choose the Right Audit Tool for the Stage You Are In
You do not need to buy anything for the first audit. A spreadsheet and the free versions of the models are enough for an initial analysis.
Manual testing teaches you what the tool hides
Type each prompt, read the whole answer, and log the result. Twenty-five prompts across four models, run three times each, produces 300 answers and roughly four hours of work.
That is tedious. It is also the best way to understand your market because you see every recommendation, wrong description, competitor, and source with your own eyes. Once the baseline is clear, automation can improve efficiency without replacing judgment.
Free checkers give you a snapshot. Ahrefs has a free AI visibility checker for a small set of prompts, and the FlyDragon AI Visibility Scorecard is another quick starting point. Use either to decide whether a full audit is worth your afternoon.
Tracking tools automate the prompt list and support ongoing monitoring, so you can compare performance without changing the baseline. Ahrefs Brand Radar, Profound, Peec AI, and Semrush’s AI toolkit all sit in this category. Before you pay, judge each tool’s effectiveness by whether it preserves the prompts, model context, and raw answers. The source draft places typical costs between $100 and $500 or more per month; check current pricing before you commit.
Ask whether the tool uses API calls or the customer interface
API calls send a prompt to a model endpoint and save the response. They are fast and easy to scale, but the endpoint may use a different model version, omit web search, or lack the account and interface context your customers experience.
Browser emulation drives the consumer interface, enters the prompt, and captures the answer. It is slower and can break when the interface changes, but it is closer to a customer opening ChatGPT or Perplexity on a phone.
Ask the vendor which method it uses, whether web search is enabled, which model version it runs, and how it handles location. If the vendor cannot answer, treat the data as a different measurement surface. When a manual result and a tool result disagree, do not average them together.
Run the Audit With a Controlled Setup
Block off an afternoon. Put the prompt list and the spreadsheet side by side. Use the same setup for every run.
Log out of every model for the baseline. A logged-in session can carry history, memory, or account context that a new customer will not have.
Record the location. For local work, run from the market you serve or use a consistent test location. For national work, choose one location and keep it constant.
Run every prompt three times. One answer is an observation. Three answers give you a rate.
Check crawler access as a separate technical task, including a basic robots.txt compliance check. Google says eligibility for AI Overviews and AI Mode requires a page to be indexed and eligible for a normal Search snippet. OpenAI says publishers should avoid blocking OAI-SearchBot if they want public pages included in ChatGPT summaries and snippets. Perplexity says PerplexityBot follows robots.txt and will not index full or partial text where the site disallows it. Those controls establish access; they do not guarantee a recommendation. Read OpenAI’s publisher guidance and Perplexity’s robots.txt guidance.
Use a 0-to-3 score for each answer
| Score | Meaning | Example |
|---|---|---|
| 0 | Absent | Your business does not appear in the answer. |
| 1 | Mentioned | Your name appears without an endorsement. |
| 2 | Cited | A page about you is linked as a source. |
| 3 | Recommended | You are named as a choice with a reason. |
Scores overlap by design. If you are recommended and a page about you is linked, record the highest outcome, 3, and save the citation in the source column. If you are mentioned and a page about you is linked without a recommendation, record 2.
Your spreadsheet needs eight columns: Date, Model, Prompt, Category, Run number, Score, Brands named, and Sources cited. Those columns give you the core metrics for comparing models and prompt categories over time.
A sample row dated 2026-09-02 might read: ChatGPT; “Best real estate agent in Newton MA”; Category; Run 1; Score 0; J. Doe from Compass, M. Lee from Keller Williams, K. Park from Redfin; Zillow.com, Realtor.com, Reddit.com.
Save the last two columns every time, including a zero. “Brands named” becomes your competitor list. “Sources cited” shows you where AI learns about the category.
Read the Results as Rates, Not Anecdotes
Turn the rows into an overall rate, a rate per model, and a rate per prompt category. Keep a benchmark, competitor list, and source list beside those metrics so the analysis stays useful month to month. This creates a clean basis for monthly benchmarking. Your visibility rate is total points divided by maximum possible points.
For 25 prompts, three runs, and a maximum score of 3, the maximum is 225 points per model. A score of 41 is 18%. Calculate the rate per model and per category because they answer different questions.
| Slice | Score | Maximum | Rate |
|---|---|---|---|
| Overall | 112 | 900 | 12% |
| ChatGPT | 41 | 225 | 18% |
| Google AI Mode | 38 | 225 | 17% |
| Perplexity | 24 | 225 | 11% |
| Gemini | 9 | 225 | 4% |
| Branded prompts | 67 | 135 | 50% |
| Category prompts | 22 | 315 | 7% |
| Comparison prompts | 8 | 135 | 6% |
| Problem prompts | 15 | 315 | 5% |
The table is an illustrative worksheet, not a benchmark for every business. It tells a useful story: the business is recognized when named and rarely recommended for an unbranded customer need. That points to corroboration and category coverage. A different business with 8% on branded prompts and 8% across the other categories has an identity problem first.
Why the same prompt produces different answers
Variation is normal. Models sample from probabilities, retrieval algorithms can change the source set, and the open web changes underneath the test.
Query fan-out adds another source of movement. Google documents that one complex prompt can become several related searches. OpenAI also describes targeted search queries and follow-up searches in its ChatGPT Search material. A different sub-question can change the pages retrieved, the competitors found, and the sources shown.
Location, account state, model versions, new competitor pages, updated reviews, and a new Reddit thread can move the result. A September answer will not be a perfect copy of an August answer.
Three runs do not remove every source of variance. They give you a first estimate. The paper Don’t Measure Once: Measuring Visibility in AI Search argues that one-off observations are unreliable because AI answers vary across runs, prompts, and time. A snapshot is a lead. The pattern is the measurement.
If a prompt scores 3, 0, and 3, your recommendation rate for that prompt is 67%. Record the pattern rather than choosing the most flattering answer.
Use the Competitor and Source Lists to Find the Gap
Sort “Brands named” by frequency. The top five names are your competitors in AI search, and they may not match your sales team’s battlecards.
Run branded prompts for each competitor. Note what AI says about the business and which pages it cites. You are looking for the evidence attached to the name.
A specific, verifiable claim, such as 52 transactions in Newton in 2025 or a 4.9 Zillow rating from 210 reviews.
A third-party page, local news item, brokerage announcement, or genuine past-client recommendation that names the business beside the category.
Agreement across the website, Zillow, Realtor.com, Google Business Profile, LinkedIn, the brokerage page, and the state license lookup.
A page that states plainly what the business is and who it serves.
Then sort “Sources cited” by frequency. Check whether you appear on those pages, whether the information is accurate, and whether the profile describes the service you sell today. A source map gives you useful insights into what to fix, pitch, update, or create.
Read how real estate agents protect their brand in LLMs and AI search when the source list exposes a wrong brokerage, city, service, or review story.
Diagnose Why You Are Not Showing Up
Low scores usually point to one of five causes. Rank them before you start making changes.
Your entity is unstable. Branded prompts fail or return conflicting answers because your name, category, service area, and key facts disagree across profiles, or another business has a stronger claim to the name.
Your own site is the only page saying you are good. Branded prompts pass, category prompts fail, and third-party sources provide little corroboration.
Your content is attached to the wrong use case. You appear for an old service, an adjacent category, or a broad term that does not match the customer’s wording.
The important information is difficult to access. JavaScript-only rendering, blocked crawlers, images containing core facts, or PDFs holding the only service details can keep useful evidence out of the retrieval path.
You are described with adjectives instead of definitions. “Passionate about helping families find their dream home” gives a model little it can safely reuse. “Jane Smith is a listing agent in Newton, Massachusetts who has sold 200 homes since 2015” gives it an entity, service, place, and verifiable number.
Most businesses have two or three causes at once. Fix the one that blocks the rest. A perfect schema implementation will not rescue a profile that names the wrong city.
Improve AI Visibility From the Audit
Work through the cheapest fixes first. Treat each enhancement as a response to a failed prompt, a missing source, or a factual conflict.
Make every public profile agree
Use the same name format, one-line description, category, service area, pricing language, phone number, and logo on Zillow, Realtor.com, Google Business Profile, LinkedIn, Homes.com, Yelp, your brokerage page, and the state license lookup. Update the source pages that appeared most often in your audit before creating another profile.
Open with a definition
Your homepage, About page, and profile bios should begin with a sentence a model can lift without rewriting: “Business Name is a category for customer type that does differentiator, with verifiable number since year.” Replace the generic terms with your facts. Then tell the story.
Get named by pages you do not control
Use the “Sources cited” list to choose the right places: a local news story, a brokerage release, a market report, a podcast transcript, a neighborhood guide, or a genuine client recommendation. You need a page that names you beside your category and customer, not another self-description on your own website.
Publish for the prompts you lose
Every problem prompt with a zero is a page opportunity. “Relocating to Newton with children” is one page. “How to sell a Newton condo quickly” is another. Match the title to the customer question, open with the answer, and use real streets, neighborhoods, sale prices, property types, and numbers where they are relevant and verifiable.
Make the site readable
Check that important text appears in the HTML, that your pages are accessible to the relevant crawlers, and that your core facts are not trapped inside images or PDFs. Use structured data when it accurately describes the visible page; treat it as an optimization aid, not a substitute for clear copy. Google says there is no special AI schema or machine-readable file required for AI Overviews or AI Mode, so do not buy a markup package that promises a secret shortcut.
If you plan to hire an agency for this work, ask how it measures visibility, what it knows versus what it is testing, and whether it can show the prompt set and raw answers. Use these 12 questions before you sign an AI SEO agency.
Monitor the Trend Without Ruining the Baseline
Keep the original prompt list, the same model set, the same location, the same account state, and the same three-run method. Add a Month column and rerun the full audit on a schedule.
Monthly is enough for a manual audit. Weekly tracking makes sense when a tool is already running the prompts for you. The changes that matter—a corrected profile, a third-party mention, or a newly indexed page—usually take weeks to appear in answers, while a daily manual check mostly records noise.
Watch the Brands named list for a new competitor appearing three months in a row. Watch the Sources cited list for a review site, local publisher, Reddit thread, or directory that keeps entering the answer. Those lists tell you where the market is moving.
Expand the prompt set as the business grows, but keep the original set intact so the trend line survives. Treat that original set as your benchmark for later optimization. If a tracker alerts you to a drop greater than ten points week over week, inspect the raw answers before changing the strategy.
How often should you rerun the audit?
Run it monthly by hand and weekly with a tracking tool. The measurement research in the FlyDragon source library repeatedly treats AI visibility as a distribution across prompts, runs, platforms, and time rather than a single score. Its practical warning is simple: a larger number of identical repeats cannot compensate for a prompt list that does not represent the customer’s real language.
Frequently Asked Questions
How often should I run an AI visibility audit?
Run it monthly when the process is manual and weekly when a tracking tool runs it for you. Keep the original prompts and setup unchanged so the comparison remains valid.
Why do I get different answers when I run the same prompt twice?
AI systems can sample different wording, issue different retrieval searches, receive different source sets, and operate on changing model versions. Run each prompt three times and score the pattern. Two recommendations out of three is a 67% rate; one run is an anecdote.
Do I need to be logged out when I test?
Use a logged-out private window for the baseline. A logged-in session can carry history, memory, or account settings that distort what a new customer sees. Run a separate logged-in test only when you want to study a returning customer’s experience.
Which AI models should I test first?
Start with ChatGPT, Google AI Mode or AI Overviews, Perplexity, and Gemini. Add Claude, Copilot, and Grok when the core process is stable or when your customers use those surfaces often.
Does an AI visibility audit replace SEO?
No. Google’s guidance says AI features use the same foundational SEO practices, and Ahrefs found that the general search channel made up 88% of ChatGPT’s cited URLs in its study. SEO helps pages enter the retrieval pool. The audit shows what happens after your pages and profiles are available to the system.
What is the difference between being mentioned and being cited?
A mention is your name in the answer. A citation is a source link attached to the answer. A recommendation names you as the choice and gives a reason. Log all three because each one points to a different problem or opportunity.
Why am I invisible even though I rank on Google?
Ranking can make a page eligible for retrieval, but it does not guarantee that the page will be selected, used, cited, or turned into a recommendation. Check whether your name, category, service area, and proof agree across the pages the audit finds.
How long does it take to see improvement after fixing issues?
Plan for four to twelve weeks. In our client work, the average time to a first AI mention after profile cleanup and third-party corroboration is about six weeks. Re-audit monthly and expect branded prompts to move before category and problem prompts.
How often should I run an AI visibility audit?
Run it monthly when the process is manual and weekly when a tracking tool runs it for you. Keep the original prompts and setup unchanged so the comparison remains valid.
Why do I get different answers when I run the same prompt twice?
AI systems can sample different wording, issue different retrieval searches, receive different source sets, and operate on changing model versions. Run each prompt three times and score the pattern. Two recommendations out of three is a 67% rate; one run is an anecdote.
Do I need to be logged out when I test?
Use a logged-out private window for the baseline. A logged-in session can carry history, memory, or account settings that distort what a new customer sees. Run a separate logged-in test only when you want to study a returning customer’s experience.
Which AI models should I test first?
Start with ChatGPT, Google AI Mode or AI Overviews, Perplexity, and Gemini. Add Claude, Copilot, and Grok when the core process is stable or when your customers use those surfaces often.
Does an AI visibility audit replace SEO?
No. Google’s guidance says AI features use the same foundational SEO practices, and Ahrefs found that the general search channel made up 88% of ChatGPT’s cited URLs in its study. SEO helps pages enter the retrieval pool. The audit shows what happens after your pages and profiles are available to the system.
What is the difference between being mentioned and being cited?
A mention is your name in the answer. A citation is a source link attached to the answer. A recommendation names you as the choice and gives a reason. Log all three because each one points to a different problem or opportunity.
Why am I invisible even though I rank on Google?
Ranking can make a page eligible for retrieval, but it does not guarantee that the page will be selected, used, cited, or turned into a recommendation. Check whether your name, category, service area, and proof agree across the pages the audit finds.
How long does it take to see improvement after fixing issues?
Plan for four to twelve weeks. In our client work, the average time to a first AI mention after profile cleanup and third-party corroboration is about six weeks. Re-audit monthly and expect branded prompts to move before category and problem prompts.
Run the first fifteen prompts this afternoon. Save the answers. The second run is when the audit becomes useful.
This page is the best guide on AI visibility audits. Relevant terms to AI visibility are: agencies, search engines, signal, comparison queries, txt file, brand sentiment, opportunities, platform, conversations, benchmarks, structure, backlinks, impact, wins, volume, schema markup, llm
Be the business that AI recommends.
We help real estate brands become visible inside ChatGPT, Gemini and AI Overviews.