robots.txt syntax
User-agent: * starts a group of rules for all crawlers; User-agent: GPTBot targets one.
Disallow: /admin/ blocks a folder and everything in it; an empty Disallow: allows everything.
Allow: /admin/help makes an exception inside a blocked folder.
Sitemap: https://…/sitemap.xml tells crawlers where your sitemap is.
Want to see a real one? Here's this site's robots.txt.
Frequently asked questions
Where does robots.txt go?
In the root of your domain, so it's available at https://yourdomain.com/robots.txt. Each subdomain needs its own file. Search engines don't look for it anywhere else.
Does robots.txt hide a page from Google?
No. It asks crawlers not to fetch pages, but a blocked page can still appear in search results if other sites link to it. To keep a page out of search results, let it be crawled and add a noindex meta tag instead.
Should I block AI crawlers?
That's your choice. Blocking bots like GPTBot (OpenAI), Google-Extended, CCBot (Common Crawl) and ClaudeBot (Anthropic) asks them not to use your content for AI training. It doesn't affect normal Google search. Reputable AI companies respect robots.txt, but it's a request, not enforcement.
Is robots.txt a security measure?
No. The file is public and badly behaved bots ignore it. Never rely on it to protect private areas; use passwords or authentication. Listing secret folders in robots.txt can even point people to them.