In this post
- What Is robots.txt and Why Does It Matter for SEO?
- How Googlebot Reads robots.txt
- robots.txt Syntax: The Basics
- What to Disallow: Common Use Cases
- Admin and Login Pages
- Faceted Navigation and Filter Parameters
- Staging and Development Environments
- Search Results and Session URLs
- What NOT to Disallow
- CSS and JavaScript Files
- Your Sitemap Directory
- Pages You Want Indexed
- The Allow vs Disallow Priority Rule
- Validating Your robots.txt
- Crawl Budget: Does robots.txt Still Matter?
- FAQ
- Does robots.txt affect Google Discover or image search?
- Can robots.txt block specific query parameters?
- How do I allow one crawler and block all others?
- Does robots.txt affect Core Web Vitals?
This post was published on 1 April 2025. Search engine behaviour and product features change, so check the current documentation before you act on it.
What Is robots.txt and Why Does It Matter for SEO?
robots.txt is a plain-text file hosted at the root of your domain (e.g. https://example.com/robots.txt) that tells search engine crawlers which URLs they are and aren't allowed to crawl. It's part of the Robots Exclusion Protocol - a voluntary standard that all major crawlers (Googlebot, Bingbot, DuckDuckGo's DuckAssistBot) honour by default.
The critical distinction: robots.txt controls crawling, not indexing. Disallowing a URL in robots.txt doesn't prevent it from being indexed if other pages link to it. For index control, use the noindex meta tag or X-Robots-Tag header instead.
How Googlebot Reads robots.txt
Googlebot fetches robots.txt before crawling any URL on your domain. If it can't access the file (server error 5xx), it will wait and retry before crawling. If the file returns 404, Google treats the entire site as unconstrained - all URLs are crawlable.
Google caches robots.txt for up to 24 hours. Changes you make today may not be reflected in Googlebot's behaviour until the next cache refresh.
robots.txt Syntax: The Basics
A robots.txt file is made up of groups, each consisting of a User-agent line followed by one or more Allow or Disallow directives:
User-agent: *
Disallow: /admin/
Disallow: /private/
Allow: /admin/public-assets/
User-agent: Googlebot
Disallow: /staging/
Sitemap: https://example.com/sitemap.xml
Key syntax rules:
User-agent: *applies to all crawlers- More specific user-agent rules override the wildcard for that crawler
Allowtakes precedence overDisallowwhen paths are equal length- Trailing slash matters:
/adminand/admin/behave differently - Comments start with
# - Always include your
Sitemap:directive at the end
What to Disallow: Common Use Cases
Admin and Login Pages
Block CMS admin panels, login pages, and dashboard routes from crawlers. They provide no SEO value and waste crawl budget:
Disallow: /wp-admin/
Disallow: /admin/
Disallow: /login
Disallow: /dashboard/
Faceted Navigation and Filter Parameters
E-commerce sites generate thousands of duplicate or near-duplicate URLs through category filters (?color=red&size=M). Block these at the robots.txt level to preserve crawl budget for canonical product pages:
Disallow: /*?color=
Disallow: /*?size=
Disallow: /*?sort=
Alternatively, handle this in Google Search Console's URL Parameters tool (legacy) or via canonical tags on each generated URL.
Staging and Development Environments
Block your entire staging domain in its own robots.txt file - don't rely on noindex alone:
User-agent: *
Disallow: /
This prevents accidental indexing of staging content and eliminates it as a duplicate content source.
Search Results and Session URLs
Disallow: /search?
Disallow: /*?session=
Disallow: /*?ref=
What NOT to Disallow
These are the most common robots.txt mistakes that harm SEO:
CSS and JavaScript Files
Blocking /wp-content/ or /assets/ prevents Googlebot from rendering your pages correctly. Google needs to access your CSS and JS to understand how your page looks and behaves. Blocking these files causes rendering failures that suppress rich results and hurt Core Web Vitals assessments.
Your Sitemap Directory
Never disallow the path where your sitemap is hosted. Always explicitly allow it if it falls under a blocked directory.
Pages You Want Indexed
This sounds obvious, but misuse of wildcard patterns (/*) routinely catches important pages. Always test your directives before deploying.
The Allow vs Disallow Priority Rule
When Googlebot finds both an Allow and Disallow that match the same URL, it uses the most specific rule. If specificity is equal, Allow wins. Example:
Disallow: /private/
Allow: /private/public-report.pdf
Googlebot will crawl /private/public-report.pdf but disallow all other /private/ URLs.
Validating Your robots.txt
Use Skymoon's Robots.txt Validator to check your file against:
- Syntax errors (invalid directives, malformed user-agent groups)
- Unintentional broad blocks (wildcards that catch more than intended)
- Missing Sitemap directive
- Disallowed paths that conflict with your sitemap URLs
- Common crawler-specific issues (Googlebot vs. Bingbot handling)
Google Search Console also has a built-in robots.txt tester under Settings → robots.txt. Use both.
Crawl Budget: Does robots.txt Still Matter?
For small sites (under 1,000 pages) with fast server response times, crawl budget is rarely a bottleneck. For large e-commerce sites, news sites, or sites with significant JavaScript rendering overhead, crawl budget optimisation via robots.txt can meaningfully improve how quickly new and updated content gets indexed.
The signals that reduce crawl budget allocation: slow server response times, large numbers of redirect chains, disallowed CSS/JS blocking rendering, and excessive 404 responses. Fix these before treating robots.txt as a crawl budget fix.
FAQ
Does robots.txt affect Google Discover or image search?
Disallowing a URL in robots.txt prevents Googlebot from crawling it, but the URL can still appear in Google Discover or image search if it's linked from other pages. To fully prevent a URL from appearing in Google's results, use the noindex meta tag on the page itself - but Googlebot must be able to crawl the page to read the tag.
Can robots.txt block specific query parameters?
Yes. Use wildcard syntax: Disallow: /*?utm_ blocks all URLs containing ?utm_. Be precise - Google's robots.txt parser treats * as a wildcard matching any character sequence.
How do I allow one crawler and block all others?
Use a specific user-agent block followed by a wildcard block:
User-agent: Googlebot
Allow: /
User-agent: *
Disallow: /
This allows only Googlebot to crawl the full site. All other crawlers are blocked entirely.
Does robots.txt affect Core Web Vitals?
Indirectly. If you block CSS or JavaScript files that are needed to render your pages, Googlebot will score your pages based on an incomplete render - potentially missing above-the-fold content or layout shift elements. This can affect your CWV data in Google Search Console even if real users see the page correctly.
Put this into practice
These products do the work described above. The first run is free, without an account.

