July 28, 2026
Short answer: AI search readiness means your content can be found, parsed, and quoted by AI assistants such as ChatGPT, Perplexity, Google’s AI Overviews, Gemini, and Claude. It rests on three things: letting the AI crawlers in, structuring content so answers can be lifted out cleanly, and being credible enough to cite. This checklist covers all three, in the order that matters.
Traditional SEO gets you ranked. AI search optimization gets you cited — and they are not the same job. A page that ranks on the second page of Google can still be quoted in an AI answer if it is well structured and credible, because AI systems select passages, not positions.
Before publishing this checklist we ran it against ourselves, using twelve queries a real buyer might ask an AI assistant when looking for what we do — including our home market.
Revolution Web was cited in zero of the twelve. Not ranked low. Absent.
What was cited, repeatedly, was Clutch and DesignRush — each appearing in five of the twelve answers. The assistants were not reading agency websites at all. They were reading directories and summarising them. That single result is why Part 5 of this checklist exists, and why we would put it above most of the on-site work if you can only do one thing.
A second finding was less obvious and more uncomfortable. An unrelated company with a similar name held the directory profiles for that name, complete with reviews. Anyone asking an assistant about us was being shown somebody else’s reputation. Being absent is not neutral — it leaves room for someone else to be mistaken for you. That is item 23 on this list, and we found it the hard way.
We also rebuilt our own site’s performance during the same period, taking Google PageSpeed Insights from 31 to 75 on mobile and 61 to 87 on desktop. It was worth doing for visitors and for rankings. It moved our citation count by nothing on its own. We mention it because speed work is regularly sold as an AI-search strategy, and on our own evidence it is not one.
As of publication we have completed the audit, the structural and schema work, the llms.txt and pricing files, and the monitoring loop. We have not finished the directory profiles — the highest impact item on the list — because they require verified signups that take real time. We will publish the re-audit when it runs, including the result if the number has not moved.
Figures recorded July 2026.
Nothing else on this list matters if the crawlers are blocked. Check these first.
Open yourdomain.com/robots.txt and check for Disallow rules targeting these user agents:
GPTBot and ChatGPT-User — OpenAI / ChatGPTPerplexityBot — PerplexityClaudeBot and anthropic-ai — Anthropic / ClaudeGoogle-Extended — Google Gemini and AI OverviewsBingbot — Microsoft CopilotBlocking any of these prevents that platform from citing you. Many sites block them by accident, often because a plugin or a well-meaning “protect my content from AI” setting added the rules. This is a genuine business decision — blocking prevents training use but also prevents citation — so make it deliberately rather than by default.
A malformed robots.txt can be interpreted far more restrictively than intended. Check it renders as plain text, uses correct directive syntax, and declares your sitemap.
Several AI crawlers do not execute JavaScript. If your main content is client-rendered, those systems see an empty page. Test by viewing the raw HTML source and confirming your actual body copy is present in it.
AI systems cannot read what sits behind a form, a login, or a paywall. Your best explanatory content should be open if you want it cited.
Place a plain-text or markdown file at yourdomain.com/llms.txt summarizing what your organization does, who it serves, and where the important pages are. It is an emerging convention rather than a formal standard, but it is inexpensive and gives AI systems a clean, unambiguous overview.
AI agents increasingly evaluate providers on a buyer’s behalf. If your relevant commercial information is only available inside a rendered page or behind “contact us”, agents skip you in favor of a competitor they can parse. A simple markdown file at /pricing.md works — and if you quote custom rather than fixed prices, say so explicitly and describe what drives a quote. Stating your model plainly is far better for an agent than silence.
AI systems quote passages, not pages. Every claim that matters should stand alone without the surrounding context.
Put the answer in the first sentence, then explain. Content that builds to a conclusion at the end gets skipped, because the extractable passage never appears.
That is about the length AI systems tend to lift as a self-contained quote. Longer answers get truncated, sometimes in ways that change the meaning.
Use “How much does a website cost?” rather than “Investment considerations.” Heading text is a strong signal for matching a query to a passage.
Comparison content is among the most-cited formats in AI answers, and tables are dramatically easier to parse than prose describing the same trade-offs. Any “X vs Y” topic should contain an actual table.
Sequential steps in a numbered list extract cleanly. The same steps as flowing paragraphs usually do not.
Use the phrasing your customers actually use, including the awkward ones. Each answer should be self-contained.
For any “what is X” query you want to win, put a clean one-or-two-sentence definition in the opening paragraph.
Paragraphs carrying three ideas cannot be quoted without carrying all three. Split them.
Research on generative engine optimization published at KDD 2024 by Aggarwal et al. tested a range of content changes and found that adding citations, statistics, and quotations produced the largest visibility gains — while keyword stuffing measurably reduced visibility. The direction of that finding is consistent with what these systems reward: verifiable specificity over repetition.
Claims traceable to a named source are more citable than assertions. Link to the original research rather than to an article summarizing it.
“Improves conversion” is unquotable. “Reduced cost per acquisition from $300 to $105 over eight weeks” is quotable. Attach a date to every statistic so its currency is visible.
An author name, a real role, and a short bio establishing relevant expertise. Anonymous content is weaker on every trust signal these systems use.
Recency is weighted heavily. Undated content loses to dated content even when it is better. Update the date only when you genuinely revise the content.
Quarterly for topics you actively compete on. Remove statements that have become outdated rather than leaving them to erode trust.
In traditional SEO, keyword stuffing is merely ineffective. In AI search it appears to be actively counterproductive. Write for a reader.
| Content type | Schema | What it enables |
|---|---|---|
| Blog posts and articles | Article / BlogPosting |
Author, date, and topic identification |
| FAQ sections | FAQPage |
Direct question-and-answer extraction |
| Step-by-step guides | HowTo |
Step extraction for process queries |
| Products | Product |
Attribute and availability extraction |
| Your business | Organization / LocalBusiness |
Entity recognition and disambiguation |
| Comparisons | ItemList |
Structured comparison data |
Use Google’s Rich Results Test and the Schema.org validator. Invalid structured data is frequently ignored entirely, which means the effort produced nothing.
If another business shares or resembles your name, explicit Organization markup with your address, founding details, and sameAs links to your verified profiles helps AI systems tell you apart. Businesses with a common or contested name are routinely conflated in AI answers, and this is the main defense.
This is the part most checklists omit, and it is often the highest-leverage item. AI systems cite where you appear, not only what you publish. For many organizations, the fastest route to being mentioned in AI answers is a third-party source rather than their own blog.
In most sectors, a small number of directories and review platforms are disproportionately cited in AI answers about “best provider” queries. Identify which ones appear in AI answers for your category and make sure your profile exists, is complete, and is accurate.
Name, address, and description should match exactly across your site, directories, and social profiles. Inconsistency creates ambiguity, and ambiguity gets resolved against you.
Community discussions and question-and-answer sites are cited meaningfully often. Contribute genuinely useful answers — promotional posting is filtered and damages the profile it comes from.
An independent article listing you among credible options carries more weight in AI answers than your own claim to be the best.
Choose 10–20 queries a genuine buyer would ask. Run each through ChatGPT, Perplexity, and Google, and record whether you are cited, who is cited instead, and which page of theirs was used. This takes an afternoon and is the only way to know where you actually stand.
Monthly or quarterly, using the same query set so results are comparable. Track share of citations against competitors rather than raw counts.
Segment traffic from AI assistants in your analytics. Volumes are typically smaller than search but the visitors tend to arrive further along in their decision.
It builds on SEO rather than replacing it. Traditional SEO — crawlability, site speed, useful content — remains the foundation. AI search optimization adds passage-level structure, explicit credibility signals, and presence on third-party sources that AI systems draw from.
Structural changes can be picked up within weeks as content is recrawled. Credibility and third-party presence take considerably longer, because they depend on other sites and platforms updating. Expect a meaningful shift over months, not days.
It is a real trade-off. Blocking reduces the use of your content for training but also removes you from citation in AI answers. A common middle position is allowing search and citation crawlers while blocking training-only crawlers such as CCBot. Decide deliberately.
They can, particularly for informational queries a user can resolve without clicking. That is precisely why citation matters: if the answer is going to be summarized regardless, being the cited source preserves visibility and referral traffic that would otherwise go entirely to a competitor.
Checking that AI crawlers are not blocked. It takes two minutes, and if they are blocked, nothing else you do can work.
Yes, and the third-party presence items matter more. AI answers to “best [service] near [place]” lean heavily on directories, review platforms, and local citations rather than on the businesses’ own websites.
Work through Part 1 today — it is short, and access problems make everything else pointless. Then run the baseline audit in item 28 so you have something to measure against. Structure and credibility work is straightforward once you know which queries you are losing and to whom.
Revolution Web runs this process for its own site and for clients, including the baseline citation audit and the structural and schema work that follows. If you would like to know where you currently stand in AI answers for your category, get in touch.
July 28, 2026
July 28, 2026
July 28, 2026