Home / SEO Tools / Robots.txt Tester
SEO Tools

Robots.txt Tester

Test robots.txt with standards-aware group selection, wildcard and percent-encoding precedence, line diagnostics, crawler matrices, 500 KiB limits and deployment fetch-behavior evidence.

RFC 9309 core · Google crawler semantics · local evidence

Robots.txt Protocol & Fetch Audit Studio

Trace exactly which crawler group and rule wins, normalize UTF-8 and percent-encoded paths, lint every directive, model Google fetch-status behavior, and export a reproducible audit before deployment.

Most-specific crawler group + merged duplicates* / terminal $ + least-restrictive tieUTF-8 / percent-encoding trace500 KiB Google processing windowHTTP 2xx/3xx/4xx/5xx preflightCSV · JSON · Print · local summaries

No proxy is used. Cross-origin CORS policy may block the fetch even when a crawler can retrieve the file.

No live fetch attempted.

Crawler product tokens

Ready.
Raw UTF-8 bytes—
Google window—
User-agent groups—
Allow / Disallow rules—
Sitemap records—
Parser findings—

Crawler access matrix

Each cell shows the local verdict and exact winning rule. The rule specificity number counts the canonicalized pattern length, including wildcard/end-anchor syntax used by Google-style precedence.

Canonical comparison trace

Line-by-line parser audit

LineFieldValueStatusEvidence

Findings

Parsed groups

GroupUser agentsRulesRule lines

Google robots.txt fetch-behavior preflight

This separate model explains why a syntactically perfect file can behave differently when the live /robots.txt request returns a redirect, 404, 429 or server error.

This is a documented-behavior explainer, not a live server monitor. A 429 is deliberately not collapsed into the ordinary 4xx “missing file” case.

Evidence handoff & repeat use

What this Studio verifies

Group selection before rule precedence

The most-specific matching crawler product token is selected first. Duplicate groups at that specificity are merged; the global * group is not mixed into a more-specific group.

Pattern length really matters

Google-style path matching supports * and terminal $. The most-specific matching rule wins and an equally specific conflict resolves to the least restrictive rule.

Raw UTF-8 and percent forms converge

The trace canonicalizes raw non-ASCII path text and common percent-encoded unreserved octets before comparison so encoded and Unicode forms can be reviewed explicitly.

robots.txt is crawl control, not access control

A blocked URL can still appear in search results without a snippet if discovered elsewhere. Sensitive content needs real authorization, not robots.txt.

Truth boundary: this page is a local RFC 9309 / Google-style preflight. It cannot impersonate Googlebot, bypass CORS, prove what a search engine fetched, validate Search Console state, guarantee crawler compliance, control indexing, or observe server redirects/status/cache headers unless the browser is allowed to fetch them. Different crawlers may implement extensions differently.

Start with the actual winning group

A robots.txt decision is not “first matching line wins.” The crawler group is selected first, duplicate groups at the same product-token specificity may be merged, and only then are matching Allow/Disallow rules compared.

Percent encoding is part of the decision

Raw UTF-8 paths and percent-encoded forms can represent the same request path. The Studio exposes the canonical comparison path so a rule does not appear to fail only because one side uses encoded octets.

Check the live fetch separately

A correct draft does not help if /robots.txt redirects too many times, returns the wrong status, or is temporarily unavailable. Use the HTTP behavior panel as a deployment preflight and verify the real response with server/Search Console evidence.

Remember the 500 KiB processing window

Google ignores robots.txt content after 500 KiB. Keep essential rules early and consolidate oversized files rather than assuming trailing directives will be processed.

Do not use robots.txt as a security boundary

Robots rules are crawler instructions, not authentication. Protect private resources with authorization and use indexing controls appropriate to the actual goal.

Search by task, tool name, or category. Press Esc to close.
Start typing to find a tool.