Search and AI
How AI search finds websites
What Google, OpenAI, and Perplexity actually say about crawl access, search visibility, and the difference between search and training.
AI search has created a market for confident-sounding advice about secret files, special markup, and guaranteed citations. The official documentation is less dramatic: make the page accessible, indexable, useful, and clear about what it says.
Google says there is no separate AI SEO requirement
Google's guidance for AI Overviews and AI Mode says the normal SEO fundamentals remain relevant and that there are no additional technical requirements or special optimizations needed to appear. A page must still be indexed and eligible to appear with a snippet in Google Search.
The same guidance recommends familiar work: allow crawling, make important content available as text, use internal links, provide a good page experience, and make structured data match the visible page. It explicitly says there is no special AI schema or machine-readable AI file required.
Search access and model training are different controls
OpenAI's crawler documentation separates robots.txt controls for OAI-SearchBot and GPTBot. OAI-SearchBot is used to surface sites in ChatGPT search, while GPTBot is used for crawling content that may be used to improve foundation models. A site can allow one and disallow the other.
Perplexity's crawler documentation makes a similar distinction. It describes PerplexityBot as the crawler used to surface and link websites in search results and says it is not used to crawl content for foundation-model training. Treat those user-agent rules as separate product decisions, not as one global AI switch.
Robots.txt manages crawling, not privacy
Google's robots.txt documentation is clear that robots.txt is mainly a way to manage crawler access and request load. It is not a security boundary, and a disallowed URL can still appear in search if other pages link to it. If content must not be indexed, use an appropriate noindex or access-control mechanism instead.
For a public blog, the practical checklist is straightforward: return a successful response, allow the public article paths to be crawled, keep private application routes out of the index, and make sure the sitemap points to canonical article URLs.
Write pages that answer a real question
The technical layer only creates eligibility. Google's people-first content guidance says its systems prioritize helpful, reliable content and asks creators whether a page provides original information, substantial coverage, useful analysis, clear authorship, and enough value that a reader would not need to search again immediately.
That is a better editorial brief than writing around a keyword. Pick a narrow question, explain the terms, show the tradeoffs, cite the material that shaped the answer, and include practical steps that a reader can apply.
Make the page easy to discover and interpret
Google recommends descriptive URLs, useful internal links, readable titles, and concise descriptions. Its link guidance also recommends descriptive anchor text rather than generic labels such as 'read more'. These details help both readers and crawlers understand how an article relates to the rest of the site.
Article structured data can help Google understand a blog post's headline, author, image, and dates, but the markup must describe the visible article. It is a classification aid, not a substitute for a useful page.
Measure what is actually available
Google says AI Overviews and AI Mode are included in the overall Web search data in Search Console. It also documents a separate generative AI performance report that is rolling out to a subset of site owners, with dimensions for pages, countries, dates, and devices.
Do not promise a universal AI ranking score. Track indexed pages, organic impressions, queries, qualified visits, and conversions first. Add the AI-specific report when the property has access, then compare the pages and topics that receive impressions instead of treating every answer-engine mention as a reliable metric.
Use trend data to test the question, not invent a hack
Google Trends can help you compare phrases such as 'AI search', 'AI Overviews', 'ChatGPT search', and 'generative engine optimization'. Use that comparison to decide which reader question deserves a full explanation, then check related searches and regional interest before choosing the final angle. Google's generative AI search guide is the useful counterweight to trend excitement because it explicitly rejects special AI files and ranking hacks.
Do not turn every spelling or wording variation into its own page. Google's current guidance warns against creating separate content for every possible query variation and says there is no ideal word count, no required chunking pattern, and no special AI file that unlocks visibility. One substantial article that answers the real question is a better bet than a cluster of near-duplicates.

