Discovery
The service learns that a URL exists through links, sitemaps, submissions or other references.
A robots rule can permit a crawler while a firewall blocks it. A successful crawl can still end without indexing, retrieval or visibility.
The service learns that a URL exists through links, sitemaps, submissions or other references.
The crawler requests the page and receives a usable response without a login, challenge or infrastructure block.
The service renders, parses and evaluates the content, directives and canonical relationships.
The service decides whether and how to retain the material for future search or retrieval.
RFC 9309 standardizes the Robots Exclusion Protocol. It defines how compliant crawlers identify groups, match paths, resolve allow and disallow rules, handle errors and cache the file.
The standard explicitly distinguishes robots rules from access authorization. Sensitive information requires authentication and server-side protection, not a disallow line.
Requests exclusion from a compliant search index; the crawler must access the page to read it.
Signals the preferred version among duplicates or near-duplicates.
Moves crawlers and users to another address and can consolidate a changed location.
Lists important canonical URLs and optional modification information.
Notifies participating engines that a URL was added, updated or removed.
Provide platform-specific inspection and diagnostic evidence rather than universal status.
A file that looks correct in the repository may be altered by the host, CDN or deployment process. Verify the public URL, response code, headers, rendered content and logs after deployment.
The standards-track specification for robots.txt parsing, matching, errors and caching.
Official guidance on SEO, crawlability, original content, local details and generative Search.
Official information about OAI-SearchBot, GPTBot, referral tracking and agent accessibility.
Official roles and access guidance for PerplexityBot and Perplexity-User.
Protocol for notifying participating search engines about added, updated and deleted URLs.
A standards-based explanation of robots.txt groups, matching, error handling, caching and why the protocol is not security or access authorization.
A source-based explanation of Google crawling, indexing, SEO, query fan-out, AI Overviews, AI Mode, structured data and site-owner controls.
An explanation of OAI-SearchBot, GPTBot, ChatGPT referrals, noindex, crawl access and accessible websites for agent interaction.
A practical guide to PerplexityBot, Perplexity-User, automatic crawling, live user-requested retrieval, WAF access and monitoring.