People also ask
In practice
The gate is blunt: a blocked crawler cannot read the site, so nothing it sees can be retrieved, quoted or cited — and most blocks are inherited accidents, copied robots templates, or bot-management rules built for scraper defence that catch AI crawlers in the same net. The decision deserves to be deliberate, because the trade is real: blocking is a licensing and cost posture (denying model builders free use of your content), while permitting it is the price of admission to the citation layer — no access, no answers, regardless of everything else the programme does.
The verification is mechanical: read robots.txt as each named crawler does, test-fetch key pages with the crawlers’ user agents, and check the bot-management layer (Cloudflare and friends have per-bot AI settings that quietly default to block). The practical configuration for a brand that wants AI visibility: allow the retrieval crawlers (PerplexityBot, the search-facing bots whose index feeds live answers), decide the training crawlers (GPTBot and friends — permitting them feeds future model memory; Common Crawl’s CCBot is the broadest training path), and keep the true bad bots blocked — the distinction is per-bot rules, not a blanket.
The audit belongs in every GEO engagement’s first week: it is cheap, it is decisive, and it occasionally explains an entire year of invisible AI presence in one robots line. What access does not do: guarantee citations — it opens the door; the content, the mentions and the pages worth quoting still have to walk through it.
See also: AI training data, GEO checklist for SaaS sites.
Related service: Brand mentions.
