On this page
AI search optimization starts with something much more basic than most of the tactics people talk about. The system has to be able to access your source, retrieve it for the right question, use it as evidence, and understand which entity that information belongs to. Only then does it make sense to worry about whether the page gets cited, the brand gets mentioned, or the company gets recommended.
There also isn't one universal position to win. ChatGPT, Google AI Overviews, AI Mode, Perplexity, and Copilot all assemble answers around a prompt, a conversation, a location, and a particular moment in time. You can improve your chances of appearing. You can't control the answer.
Ranking in AI Search Has More Than One Outcome
I run AI-search campaigns at FlyDragon, mostly for real estate businesses, and I keep seeing the same mistake.
"Ranking in AI search" gets used as though it describes one thing. It doesn't.
A page can be retrieved, used, cited, mentioned, recommended, or visited. Sometimes several of those happen at once. Sometimes they don't.
Traditional search makes this easier to see because you can point to a URL sitting in a particular position. AI answers are messier. A system might read twenty pages, cite four, mention three companies, and recommend one. You might see traffic from a cited page while another page on your site contributed a fact to the answer without ever receiving a visible link.
| Outcome | What happened | What you can usually observe |
|---|---|---|
| Retrieved or fetched | The page entered the candidate source pool. | Partial evidence from server logs or platform tools. |
| Used | The answer drew language, structure, evidence, or facts from the page. | Often an inference from close source comparison. |
| Cited | A visible source link points to the page. | The citation and its surrounding claim. |
| Mentioned | The brand, product, or person appears by name. | The captured answer. |
| Recommended | The entity is selected for the user's task and given a reason. | The wording, order, conditions, and competitors named. |
| Visited or converted | The user continues to the site or business. | Analytics, calls, forms, sales, or another conversion record. |
For reporting, I keep citation rate, mention rate, recommendation rate, and referral conversions separate. You can roll them into one AI visibility score if you want a simple dashboard number, but the trade-off is obvious: once everything is blended together, you lose the explanation for why performance moved.
I expect serious AI-search reporting to start splitting these outcomes out by default over the next twelve months. A single score is useful for a quick glance. It's pretty poor for deciding what to fix.
How AI Search Systems Find and Choose Sources
Different AI-search products use different systems, so there isn't one pipeline we can point to and say, "this is how all of them work." A useful working model is still possible though.
A request gets interpreted. Searches may be issued or rewritten. Candidate sources are retrieved, merged, filtered, or reranked for the current context. Supporting passages are pulled out. The answer is generated, and citations are attached where the product provides them.
I use that sequence as a diagnostic model, not as a claim about a published algorithm.
Google documents query fan-out for its AI features, where a complex request can produce several related searches across subtopics and different data sources. A Google patent for stateful chat describes synthetic queries being built from the user's request and earlier conversation. Another filing describes initial passage retrieval followed by cross-encoder reranking and grounding-span selection.
Patents are useful because they expose possible architectures. They do not prove that a live product uses every step or signal described in them today.
What they do help show is that retrieval and final source selection can happen at different stages. That distinction matters a lot more than people think.
A blocked URL can lose before retrieval. A vague page can be accessible and indexed but still fail to match the search being generated. Another page might match perfectly and then lose during reranking because its evidence isn't strong enough. A company can even earn citations and still fail to get recommended because the broader evidence around that entity is weak.
Find the first stage where the source drops out. That's usually where the work belongs.
Start With a Query Network, Not One AI Keyword
Repeating the same phrase across a title, a few headings, and the body is a very shallow way to think about relevance.
A page becomes useful for a task when it covers the main question and the connected questions that have to be answered before the system can finish the job.
Take:
"What is the best project-management software for a remote design team?"
A decent answer may require information about design approvals, ten-person pricing, Figma integrations, Slack integrations, security, customer reviews, and direct product comparisons. The original query is one part of the problem. Those supporting searches are the rest of it.
That's the query network.
Build it in a spreadsheet with five columns:
- the question or likely search;
- the reader's intent;
- the entity and attribute being requested;
- the best format for the answer;
- the URL that owns it.
The trick is deciding where the intent changes.
One page can own the representative task and the follow-up questions required to answer it. A full comparison, pricing calculator, or implementation manual probably deserves its own URL because the job has changed.
That stops you falling into either extreme: one thin URL for every wording variation, or one gigantic guide trying to cover an entire subject whether the sections belong together or not.
Make the Right Page Eligible to Be Found
Before you get into content quality or citations, the relevant search system has to be able to access the canonical page, interpret its main content, and include it in whatever source pool that product uses.
That's only eligibility. It doesn't mean the page will be selected.
I would check the technical path in this order:
- Return a successful status for the canonical URL and redirect duplicate versions.
- Remove accidental noindex, nosnippet, robots, authentication, or firewall blocks.
- Render the main answer in accessible HTML rather than hiding it inside an image or an interaction that never loads for a crawler.
- Use a self-referencing canonical and include the URL in an accurate XML sitemap.
- Link to the page from a relevant hub or guide that search systems already know.
- Inspect index reports, server logs, and crawler documentation for the surface you care about.
Google says a page needs to be indexed and eligible to show a normal Search snippet before it can appear as a supporting link in AI Overviews or AI Mode. Its newer generative AI search guide also makes a few things very clear: Google doesn't need special AI markup, doesn't use llms.txt for Search, and doesn't require publishers to chop articles into tiny artificial chunks.
OpenAI's publisher guidance gives its crawlers different jobs. OAI-SearchBot is used for ChatGPT search visibility. GPTBot controls potential model-training use. They're separate controls and should be treated that way.
Perplexity documents a similar distinction between PerplexityBot and Perplexity-User.
I wouldn't rely on a universal "AI crawler checklist" for this stuff. Read the documentation for the product you care about. These controls change too often.
Give Every URL One Clear Job
I want to be able to look at the top of a page and answer three questions without doing much work.
What entity is this about?
What job is this page supposed to complete?
Why should I trust this particular source on that subject?
For this article, the central entity is AI-search visibility. The page is meant to teach a site owner how to improve it. The source context comes from FlyDragon's work with semantic SEO, retrieval research, and measured AI-search campaigns.
That's enough.
A 1,500-word detour into the history of large language models might technically be related to the subject, but it would make this page worse at its actual job.
This is where the language you use starts to matter as well.
"OAI-SearchBot supports search discovery" gives us an entity, an action, and a fairly clear purpose.
"OAI-SearchBot is important for AI SEO" doesn't tell us much. "Important" is doing all the work, and it can't really be checked or tied to a specific question without extra interpretation.
I also pay attention to the order of the information. Explain what the outcome is before explaining the mechanism. Explain how the mechanism works before prescribing changes to the page. Once the page work is clear, measurement and diagnosis make more sense.
The page should feel like one argument developing, not a pile of individually optimized sections.
Google's guidance also pushes against creating a fresh page for every long-tail variation. Search systems can connect synonyms and related meanings. Cover the question chain properly, and create another URL when the user is trying to do something different.
Publish Information Worth Retrieving
The best reason for a system to use your page is that the page contains something it needs.
That could be a fact, a method, a comparison, original data, a controlled test, or a useful firsthand observation that isn't available in every other summary on the web.
Formatting makes good information easier to move around. It doesn't make weak information useful.
The original Generative Engine Optimization study reported visibility gains of up to 40% for some methods inside its controlled benchmark. That's interesting research, but I wouldn't turn it into a promise that applying the same techniques today will give you a 40% lift across ChatGPT, Google, or Perplexity. The experiment doesn't support that claim.
A 2026 paper on citation selection and citation absorption looked at 602 controlled prompts, 21,143 valid citations, 18,151 fetched pages, and 72 page features. Pages that had more influence on generated answers tended to be longer, better structured, more semantically matched to the request, and richer in things like definitions, numbers, comparisons, and steps.
Again, correlation inside a designed study isn't a public weighting formula. I treat those features as clues about what makes a source useful enough to retrieve and absorb.
My read from all of this is that structure starts paying off once there's something worth structuring.
When you're adding evidence, give it enough context that someone else can understand what the number or claim means. Name the source. Include the date or method where it matters. Include the limitation when leaving it out would change the interpretation.
Roughly, I would value the evidence like this:
- First-party data with the collection method, date, and sample.
- A controlled test or case study with a baseline, intervention, and result.
- Named experience from a person whose role and scope are visible.
- Primary documents, standards, and platform guidance.
- A defensible synthesis that connects those sources in a new way.
A statistic roundup is useful when the page exists to collect statistics. We already maintain FlyDragon's current AI SEO statistics for that.
On an implementation page, I use a simpler test: does this number change what the reader should do? If it doesn't, I probably don't need it.
Make Each Answer Easy to Extract
I don't think every page needs to be broken into fifty miniature answers for AI.
What matters is whether a section answers the heading and whether that answer still makes sense when it is pulled away from the rest of the article.
Usually that means answering the heading fairly early, then choosing a format that suits the information. Processes work well as numbered steps because order matters. Repeated comparisons usually belong in a table. A set can use bullets. Explanations normally read better as prose.
The content decides the format.
Definitions should keep the conditions that make them true. Processes should retain their order. Comparisons should use consistent fields. If you're making a data claim, keep the source, date, method, or limitation close enough to the number that someone can't accidentally lift the exciting part and leave the caveat behind.
Basic publishing choices help here too. Use headings that describe the question or decision. Give tables actual column headers. Write anchor text that tells the reader where they're going. Keep the important answer in visible HTML.
Just don't over-engineer it.
A page made from sentence fragments, hundreds of micro-sections, forced FAQ blocks, and the same phrase repeated over and over might look "optimized" in a spreadsheet. It usually reads terribly.
Google says there is no ideal page length and no requirement for artificial chunking. Three thousand useful words are better than five thousand words written to hit some arbitrary content target.
Build Corroboration Beyond Your Website
Your own website tells the system what you say about yourself. Independent sources help establish whether those claims line up with the rest of the web.
That can cover the company, person, product, location, reputation, category, services, or whatever else the system needs to understand.
I want basic entity facts to be boringly consistent across good sources. Names, locations, descriptions, authors, products, categories. If an important third-party profile has the wrong location or describes a service you stopped offering three years ago, fix it.
Beyond that, real corroboration comes from things like editorial coverage, expert quotes, customer reviews, good reference listings, and other evidence you've earned.
This starts mattering even more when the query moves from "tell me about this" to "which one should I choose?"
Your own guide might be a perfectly good source for explaining how project-management software works. If the user wants the best option for a regulated enterprise buyer, the answer may also need security documentation, independent reviews, customer evidence, and accurate product information before it can confidently recommend anything.
Fake reviews, paid forum spam, mass-produced mentions, and invented consensus are a different thing entirely. They leave a dirty evidence trail, and Google's current guidance already tells site owners to ignore inauthentic mention schemes.
Build evidence you'd be happy to show somebody.
Use Internal Links as Semantic Bridges
I mostly think about internal links as a way to explain relationships.
They help people find the next useful page, but they also show search systems how one topic or task connects to another.
For this subject, a sensible structure might look like:
AI SEO hub → ranking implementation guide → visibility audit, statistics evidence, and platform update pages.
Supporting pages can link back to the parent process where the relationship is useful. The anchor should describe what the destination page does. It doesn't need to repeat the same exact keyword every time.
Context is much more useful than hitting some arbitrary internal-link quota.
If I'm talking about measurement and link to a guide explaining how to measure AI visibility, the reason for the link is obvious. If I dump fifty vaguely related articles into a footer, I've technically created more internal links without explaining much of anything.
I also remove self-links, duplicate anchors where they serve no purpose, and links added purely because somebody thinks they need to pass "SEO juice."
Google, ChatGPT, Perplexity, and Bing Expose Different Controls
Most of the durable page work overlaps across the major AI-search products. The product-specific stuff changes much faster.
Access controls change. Source indexes are different. Citation interfaces are different. Publisher reporting is different.
So I keep that layer separate from the main methodology.
| Surface | Documented access or eligibility | First place to look | Measurement available to publishers |
|---|---|---|---|
| Google AI Overviews and AI Mode | Google Search indexing and snippet eligibility. | Technical SEO, page usefulness, and relevant internal links. | Google Search Console's generative AI reporting where available. |
| ChatGPT Search | OAI-SearchBot for search discovery; GPTBot is a separate training control. | Intended crawler access and source-backed pages that web search can retrieve. | Referral analytics plus repeated prompt tracking. |
| Perplexity | PerplexityBot for indexing and Perplexity-User for user-requested fetches. | Robots and firewall access, stable canonical URLs, and useful evidence. | Captured answers, citations, referrals, and server evidence. |
| Bing and the Copilot ecosystem | Bing crawl and index systems. | Bing indexability and page-level citation evidence. | Bing Webmaster Tools AI Performance reports citations, cited pages, and grounding queries. |
Bing says its AI Performance report does not tell publishers a citation's importance, placement, or rank. That's an important limitation. Seeing ten citations tells you your pages made it into ten answers. It doesn't tell you whether those answers leaned heavily on your source, mentioned it in passing, or endorsed the business.
For the Google-specific changes, we keep a separate review of Google's latest AI-search guidance. The implementation system on this page doesn't need rewriting every time one platform changes a report or crawler rule.
Measure AI Visibility as a Rate
A single AI answer is one observation. I've seen people put far too much weight on one run.
If you want useful measurement, repeat a controlled set of prompts across products, runs, and dates, then record the outcomes separately.
For every run, save the prompt, platform, product surface, account state, location, date, answer, cited URLs, brands named, and recommendation order.
If you change something on the site, keep the original prompt set intact. Add new questions to a second cohort rather than changing the test halfway through and then comparing the results as though the baseline stayed the same.
The basic citation-rate calculation is:
Citation rate = eligible answer runs with at least one citation to the tracked domain ÷ all eligible answer runs × 100
The paper Don't Measure Once documents variation across repeated runs, prompt wording, and time, which is exactly why I don't like one-off screenshots being treated as measurement.
It doesn't give us one magic sample size either. The number of runs you need depends on how much variance you're seeing, what decision you're trying to make, and how much it costs to collect another answer.
We've got a full process for a repeatable AI visibility audit if you want the prompt design, scoring sheet, and comparison method.
For the purpose of this guide, keep the experiment boring: compare the same things under the same conditions and save the raw answers.
Diagnose the Earliest Failed Stage
This is probably the most useful distinction in the whole process.
A crawl problem needs a different fix from a retrieval problem. A retrieval problem needs a different fix from a citation-selection problem. Being cited and never recommended is another problem again.
| Symptom | Likely stage | First evidence to inspect | First repair |
|---|---|---|---|
| The page is absent from the relevant search index. | Access or eligibility. | Index report, robots, canonical, response, and rendering. | Remove the specific technical blocker. |
| The page is indexed but never appears for the topic. | Demand, relevance, or coverage. | Prompt set, competing sources, page intent, and entity relationships. | Repair intent mapping and the missing question chain. |
| The page is fetched or surfaced but receives no citation. | Reranking, evidence, or extraction. | The competing passages and their supporting proof. | Strengthen the exact answer and the evidence behind it. |
| The page is cited but contributes little to the answer. | Use or absorption. | Distinctive facts, language, and structure carried into the answer. | Add material the synthesis needs and can attribute. |
| The brand is mentioned but never recommended. | Entity fit or corroboration. | Third-party sources, reviews, category fit, and proposition clarity. | Improve verifiable off-site evidence. |
| The result changes sharply between runs. | Context or measurement. | Prompt wording, prior conversation, account, location, product, and date. | Repeat controlled runs and report the distribution. |
The easy mistake is letting each team fix the problem they already know how to fix.
Writers rewrite copy. Developers check robots.txt. PR teams go looking for more mentions.
Sometimes that's the right answer. Sometimes it has nothing to do with the stage that's failing.
Follow the evidence first.
Your First 30 Days of AI-Search Work
Thirty days is plenty of time to fix the obvious eligibility and intent problems, publish at least one useful evidence-led improvement, and establish a baseline you can repeat.
It is not a promise that every search product will recrawl the page, select it, and start citing it inside thirty days. Those are different things.
- During days 1–5, define the tracked questions, record current outcomes, assign each intent to one URL, and find pages competing with each other.
- During days 6–10, repair indexing, crawler access, canonicals, rendering, sitemaps, and internal discovery.
- During days 11–20, tighten the central entity and source context, close the necessary question chain, add original evidence, and use the right format for each answer.
- During days 21–30, correct third-party facts, pursue legitimate corroboration, repeat the baseline conditions, and log which stage moved.
If you're hiring somebody else to do this, use these questions to ask an AI SEO agency before buying guaranteed citations or a mystery visibility score nobody can explain.
Pages rarely win because somebody discovered one magic markup trick.
They answer a real question, survive the stages that happen before synthesis, and give the system better evidence than the alternatives. Then you measure what happened and work on whichever part broke first.
Frequently Asked Questions
Can you guarantee a ranking in AI search?
No. You can improve access, relevance, evidence, extraction, and corroboration. The search product still controls its own retrieval, generation, and citations.
If somebody guarantees a specific AI ranking or citation outcome, I'd want to see very strong evidence for how they're making that promise.
Do I need llms.txt to rank in AI search?
Google says it does not use llms.txt for Search, including its generative AI features.
Another service may decide to use the file, so check its documentation. I wouldn't treat llms.txt as a universal AI-search requirement.
Does schema markup make a page rank in ChatGPT or AI Overviews?
There is no documented source supporting that guarantee.
Use structured data where it accurately describes the entity or rich-result type it was designed for. It doesn't replace visible content, crawler access, relevance, or evidence.
Can a page appear in an AI answer without ranking number one on Google?
Yes.
AI features can retrieve and rerank sources for different parts of the question, so a page doesn't need to hold a universal number-one Google position before it can appear in an AI answer.
That doesn't make normal search visibility irrelevant. Search-grounded products still rely heavily on crawlable, indexable sources.
How long does it take to rank in AI search?
There isn't one honest universal timeline.
Crawl and index refreshes, competition, the quality of the source, product behavior, off-site corroboration, and how often you're measuring can all change how quickly a result becomes visible.
Is AI-search optimization different from SEO?
I see it as an extension of SEO rather than a completely separate discipline.
The technical SEO, useful content, internal linking, and authority work still matter. AI search optimization adds another layer around query fan-out, retrieval, passage selection, answer synthesis, citations, mentions, recommendations, and repeated measurement.
How can I tell whether an AI answer used my page without citing it?
Sometimes you can't prove it.
Keep going: the AI search playbook
- LLM SEO, explained in full
- What AI-powered search is
- The best AI search engines for leads
- Writing a bio AI search can use
- How AI search changed the way people find providers
- Using video to rank in AI search
- Generating leads from ChatGPT
- What the Zillow-ChatGPT partnership means
- ChatGPT ads, explained
- Will AI replace real estate agents?
- How generative AI is changing real estate
- Why AI content is the new spam
And if you'd rather someone did this work for you — that's the job. Our GEO service exists because most businesses have the evidence and nobody's ever assembled it.
Be the business that AI recommends.
We help real estate brands become visible inside ChatGPT, Gemini and AI Overviews.