Technical Checklist for AI Crawlers and AI Search

You cannot earn AI search citations from pages bots never fetch. This checklist covers the technical basics we verify before content and authority work—part of every AI SEO engagement.
Retrieval vs training
Different user-agents exist for training corpora versus live search/retrieval. Policies change; always verify current bot documentation before blocking or allowing.
Rule of thumb:
- If you want to be found and cited in a product’s search mode, do not casually block its retrieval crawler
- If you want to limit training use of your content, that is a separate decision—don’t confuse the two
Checklist
Indexation hygiene
- Important URLs return 200 (not soft 404s)
- Self-referencing canonicals on unique pages
- HTTPS + consistent host (www vs apex)
- XML sitemap lists only canonical, indexable URLs
-
robots.txtallows what you intend to market
Page quality for extraction
- Primary content in HTML (not locked behind client-only shells)
- One clear H1; logical heading hierarchy
- Meaningful titles and meta descriptions
- Relevant JSON-LD (Organization, Service, FAQ, Article) where accurate
Performance
- Reasonable LCP on mobile
- Images compressed with modern formats
- No unnecessary third-party scripts on money pages
Tie it to strategy
Technical readiness is table stakes. Combine it with answer-first content and citation measurement.
Need a full crawl-to-citation roadmap? Inquire.


