Skip to main content

Technical SEO

Build, Check and Test Your robots.txt

Start from a preset or paste the file you have. See what each of 14 crawlers, including the AI ones, can reach, which rule decides it, and what to fix before you publish.

  • Free, no sign-up, no credits
  • 14 crawlers, including AI bots
  • Explains which rule matched
What should the file do?

Leave it empty and the file shows a TODO where it goes.

AI crawlers

Choose a rule for each, or leave it as no rule. Blocking AI search crawlers can stop AI assistants from citing your pages. Blocking training crawlers does not.

  • AI training

    Collects content that may be used to train OpenAI models. Blocking it keeps your content out of training.

  • AI search

    Finds pages to show in ChatGPT search. Blocking it can stop ChatGPT search from citing you.

  • Opened on request

    Opens pages when a person asks ChatGPT to. OpenAI notes robots.txt may not apply to it.

  • AI training

    Collects content that may contribute to training Claude models. Blocking it keeps your content out of training.

  • AI search

    Improves the quality of Claude's search results. Blocking it can reduce how often Claude cites you.

  • Opened on request

    Opens pages when a person asks Claude to. Blocking it stops Claude from reading your pages on request.

  • AI training

    Controls whether Google may use your content for Gemini training and grounding. It does not affect Google Search.

  • AI search

    Finds pages to show in Perplexity search results. It is not used for training. Blocking it can stop Perplexity from citing you.

  • Opened on request

    Opens pages when a person asks Perplexity to. Perplexity says it generally ignores robots.txt.

robots.txt
1 field still empty: Sitemap address
User-agent: *
Allow: /

# TODO: add your sitemap address, for example https://yoursite.com/sitemap.xml

Checks

1 warning to review
  • Warning

    The file does not point to a sitemap.

    Add a line such as Sitemap: https://yoursite.com/sitemap.xml.

  • Passed

    Every line is a rule search engines can read.

Test an address

See which crawlers may fetch it, and which rule decides.

Enter an address to test it.

Start from the file on your site

Fetch the robots.txt your site serves now, check it, and see what changes if you switch to the file you are building. Only your site's own /robots.txt is read. You need to be signed in, and you get 20 imports a day.

Check crawl and indexing across your whole site

Find pages blocked by robots.txt and pages at risk of dropping out of search.

Check across my site

What Is a robots.txt File?

A robots.txt file sits at the root of your site and tells crawlers which paths they may request. It is how you keep a staging site, a cart or an admin area out of search, and how you decide which AI crawlers may read your content.

It is also easy to get wrong. One stray line can block your whole site, a rule for one bot can be ignored because another group matches first, and the file only works at the root of the domain. This product builds the file from choices, then checks it the way crawlers read it, so you see the result before it goes live.

  • 14crawlers tested for every address
  • 9of them AI crawlers, each with its own control
  • 4starting points: allow all, block folders, staging and your own

How It Works

4 steps, and you stay in control of the result.

  1. Choose a starting point

    Allow everything, block some folders, a staging site that blocks everything, or write it yourself.

  2. Set the AI crawler controls

    Allow or block each AI crawler separately, with a plain note on what each one does.

  3. Test addresses

    Enter a path and see, for every crawler, whether it is allowed and which rule decided it.

  4. Fix what the checks flag

    Read the warnings, such as blocking your whole site or your scripts, then copy or download the file and upload it to your site's root.

What You Get

An address tester that explains itself

It does not just say blocked. It names the rule that matched and the group it came from, so you can see why.

AI crawler controls

Separate switches for search crawlers and AI training crawlers, with a note that blocking an AI search crawler can stop it citing you.

Risk checks

Flags a file that blocks the whole site, blocks styles or scripts, has no sitemap line, or contradicts itself.

Import your current file

Signed in, fetch the file your site serves now, check it and see a line-by-line comparison with the new one.

No guesswork in presets

Block-folders asks which folders to block. It does not assume your CMS or hide a login path you may need.

Copy or download

The file is ready to paste. A reminder says it must be at the root, for example example.com/robots.txt.

What You Get

You give it: The preset "Block some folders" with /admin/ and /cart/ blocked, GPTBot blocked and a sitemap address.

User-agent: *
Disallow: /admin/
Disallow: /cart/

User-agent: GPTBot
Disallow: /

Sitemap: https://example.com/sitemap.xml

An illustration. The file you build uses your own folders and sitemap address.

robots.txt Product vs Writing It by Hand

The syntax is short, which is why mistakes get through. The rules for which line wins are the part people forget.

By hand compared with Robots.txt Validator & Generator
TaskBy handWith this product
Writing the fileRemember the syntax and the pathsChoose a preset and your folders
AI crawlersLook up each bot's name and what it doesA switch per crawler, with a plain description
Knowing what is blockedReason through groups and wildcards in your headA table of allowed or blocked for 14 crawlers with the matching rule
Catching a disasterFind out when your pages leave searchA warning before you publish a file that blocks the whole site
Updating an existing fileEdit and hopeImport it, check it and compare before and after

Where doing it by hand goes wrong

A staging file reaches the live site

A file that blocks everything is correct for staging and a disaster in production. It happens at every launch.

Blocked styles and scripts

If a crawler cannot load your CSS and scripts it cannot render the page properly, and rankings can suffer.

The wrong rule wins

A broad Disallow and a narrow Allow interact by specificity, not by order. Reading it by eye is where people go wrong.

AI crawlers are an afterthought

Blocking the wrong AI crawler removes you from AI search answers; allowing the wrong one lets content be used for training.

Check Your robots.txt Before Crawlers Do

Build a file or paste yours and see what every crawler can reach.

What It Does Not Do, and Its Limits

So you know what to expect before you start.

  • Rules are evaluated the way Google documents them. Other crawlers can read robots.txt differently.
  • robots.txt controls crawling. It is not a way to keep a page out of search or to protect private content.
  • The tester checks a file you build or paste. It does not check that your server serves the file correctly, except through the signed-in import of your live /robots.txt.
  • Importing your live file needs a free account and is limited to 20 imports a day.
  • Some fetchers, such as those that open a page because a person asked, may ignore robots.txt.

Who Uses It

  • SEO managersCheck a client's file in a minute and show exactly which line is the problem.
  • DevelopersGenerate a correct file for a launch and test the paths that matter before deploy.
  • Site ownersDecide which AI crawlers may read your content, without learning the syntax.
  • AgenciesCompare an existing file with a proposed one and hand over a clear change list.

robots.txt Rules Worth Knowing

These come from Google's published specification and each operator's crawler documentation.

It must be at the root
The file only works at the root of a host, for example example.com/robots.txt. A file in a folder is ignored.
It controls crawling, not indexing
A blocked address can still be indexed if other pages link to it. To keep a page out of search, use a noindex tag and allow the crawl.
The most specific rule wins
The longest matching path decides. When an Allow and a Disallow match equally, Allow wins.
A bot follows one group
A crawler uses the most specific group that names it, and ignores the others. A rule for * does not add to a group that names the bot.
Do not block what pages need
Blocking CSS, JavaScript or images stops crawlers rendering the page.
Some fetchers may not obey it
Operators say certain fetchers, such as those that open a page when a person asks, may not follow robots.txt. Use real access control for private content.

Where the Rules Come From

The checks follow these official documents. Read them, and check your result with the ones marked Test.

With a free account

Check It Against Your Real Site

Signed in, you can fetch the robots.txt your site serves right now, check it and compare it with the new one. The site you pick is remembered for the next product.

Create a free account

  • Import the live file and see a before-and-after comparison
  • Your sitemap address suggested from your domain
  • Pick a site once and every product uses it

What Is New

  1. Import your live robots.txt and see a line-by-line comparison with the file you are building.
  2. New address tester that names the winning rule for every crawler, with controls for AI crawlers.

Questions, Answered

What is a robots.txt file?

A plain text file at the root of your site that tells crawlers which paths they may request. It controls crawling. It is not a way to hide private content.

How do I block AI crawlers?

Add a group for each AI crawler with Disallow: /. The product has a switch for each one. Note that blocking an AI search crawler, such as OAI-SearchBot, can stop that service citing your pages.

Does robots.txt stop a page being indexed?

Not by itself. A blocked address can still be indexed if other pages link to it. To keep a page out of search, allow crawling and add a noindex directive.

Where do I put the file?

At the root of the host, for example example.com/robots.txt. Each subdomain needs its own file.

What happens if I have no robots.txt?

Crawlers treat a missing file as permission to crawl everything. A server error is treated more cautiously, so keep the file available.

Which rule wins when two match?

The most specific one, which is the longest matching path. If an Allow and a Disallow match equally, Allow wins.

Is it really free?

Yes. Building, checking and testing use no credits and need no account. Importing your live file needs a free account.

Page checked against the product on 4 October 2026.

Built by an Agency That Uses It

Skymoon SEO Suite is made and maintained by Skymoon Infotech, a digital growth agency founded in 2023 in Ahmedabad. Each product started as a task the team repeated on client work, and is used there before it is released here.

Found a problem or have a question? Tell us. How we handle what you upload is in the privacy policy, and About says more about who we are.