webtrajans
en

Robots.txt Tester

Fetch a site’s robots.txt and test whether any path is allowed or blocked for a specific crawler, with the exact matching rule highlighted.

Sitemaps

    Rule match

      The address you enter is only used for the check and is not stored.

      How to use

      1. 1Enter the domain; we fetch its /robots.txt.
      2. 2Pick a user agent (Googlebot, Bingbot, GPTBot, * …) and enter a path to test.
      3. 3See Allowed or Blocked and which rule matched.
      4. 4Edit the file in the box to test changes before you publish them.

      How Google evaluates robots.txt

      A crawler uses the most specific group that names it (e.g. Googlebot) and ignores the * group if one does. Within that group, the rule with the longest matching path wins; if an Allow and a Disallow match with equal length, Allow wins. * matches any sequence of characters and $ anchors the end of the URL, so Disallow: /*.pdf$ blocks all PDFs. Paths are case-sensitive. This tester applies exactly these rules (as standardized in RFC 9309).

      Blocking crawling is not blocking indexing

      Disallow stops crawling, not indexing. A blocked URL can still appear in Google results — without a description — if other pages link to it. To keep a page out of Google, allow crawling and use a noindex meta tag or X-Robots-Tag header instead. And never block CSS or JavaScript files Google needs to render the page.

      Common robots.txt mistakes

      A leftover Disallow: / from a staging site (it blocks everything); blocking /wp-admin/ without allowing /wp-admin/admin-ajax.php; a robots.txt served with a 5xx error (Google may then pause crawling the whole site); using unsupported directives like noindex or crawl-delay for Googlebot (Google ignores both); and forgetting that each subdomain and protocol needs its own robots.txt.

      Frequently asked questions

      Where must robots.txt be located?

      At the root of the host: https://example.com/robots.txt. A file in a subfolder is ignored, and www and non-www are treated as separate hosts.

      How do I block AI crawlers like GPTBot?

      Add a group such as “User-agent: GPTBot” followed by “Disallow: /”. Repeat for ClaudeBot, CCBot, Google-Extended and others. Our robots.txt generator has a preset for this.

      Do edits in the box change my live file?

      No. Edits are only for testing in your browser. Upload the updated file to your server to apply them.

      How quickly does Google pick up robots.txt changes?

      Google generally caches robots.txt for up to 24 hours.

      Not happy with the results?

      Talk to Webin Agency about fast, SEO-friendly websites, e-commerce and Google Ads management.

      Get free advice

      Related tools