RFC 9309: Robots Exclusion Protocol
The standards-track specification for robots.txt parsing, matching, errors and caching.
RFC 9309 defines how compliant crawlers interpret robots.txt. It is a crawl-coordination standard, not access control, authentication or a ranking mechanism.
The Robots Exclusion Protocol uses user-agent groups and allow or disallow path rules. RFC 9309 defines how crawlers select matching groups, compare paths, resolve equivalent rules and handle the robots file under different server responses.
Standardization reduces ambiguity among compliant crawlers. It does not require every automated client to obey the file.
The RFC states that robots rules are not a form of access authorization. The file is public and the blocked path can still be visible within it. Sensitive material requires authentication, permissions and server-side controls.
This is one of the most important technical distinctions on the site: crawl preference and security policy are not interchangeable.
The protocol distinguishes successful retrieval, redirects, unavailable resources and server errors. Operational teams should understand those differences rather than assuming a missing robots file and a 5xx failure communicate the same instruction.
The standards-track specification for robots.txt parsing, matching, errors and caching.
A technical guide to discovery, crawler access, robots.txt, infrastructure blocks, directives, canonical URLs, sitemaps and indexing.
A combined editorial and technical playbook for identity, evidence, substantial pages, retrieval-ready sections, crawling, structured data, agents and monitoring.
An explanation of OAI-SearchBot, GPTBot, ChatGPT referrals, noindex, crawl access and accessible websites for agent interaction.
A practical guide to PerplexityBot, Perplexity-User, automatic crawling, live user-requested retrieval, WAF access and monitoring.