Robots.txt for SEO: Complete Guide, Examples & Best Practices
Robots.txt is a text file that gives supported web crawlers instructions about which parts of a website they may or may not crawl. Used correctly, it can help manage crawler access. Used incorrectly, it can prevent search engines from accessing important website resources.
Quick answer: A robots.txt file primarily controls crawling, not indexing. You can use directives such as User-agent, Disallow and Allow to manage crawler access to specific URL paths. Do not use robots.txt as a security mechanism or assume that blocking crawling automatically removes a URL from search results.
What Is Robots.txt?
Robots.txt is a plain-text file located at the root of a website host. It contains rules intended for web crawlers that support the Robots Exclusion Protocol.
For a website such as:
the robots.txt file would normally be available at:
The file can contain crawler-specific instructions and may also reference XML sitemap locations.
What Does Robots.txt Do?
Its primary SEO function is controlling crawler access to URL paths.
Manage Crawling
Robots.txt can tell supported crawlers not to request particular paths.
Reduce Unnecessary Crawling
On certain large or technically complex sites, it can help reduce crawling of URL patterns that provide little search value.
Reference XML Sitemaps
A robots.txt file can include a Sitemap directive pointing crawlers toward an XML sitemap.
How Robots.txt Works
A crawler can request the robots.txt file before crawling the website and interpret rules relevant to its user-agent.
Crawler → robots.txt → Applicable Rules → Crawl Allowed or Disallowed
Different crawlers may interpret supported rules according to the protocol and their own documented behavior.
Basic Robots.txt Syntax
A simple robots.txt file can look like this:
This example targets crawlers represented by the wildcard user-agent and disallows crawling of the specified path.
Understanding User-agent
The User-agent directive identifies which crawler a group of rules applies to.
The asterisk is commonly used as a wildcard for crawlers not covered by a more specific matching group.
Crawler-specific rules can also be created when there is a legitimate reason to manage different crawlers differently.
Understanding Disallow
The Disallow directive specifies a path that matching crawlers should not crawl.
This tells matching crawlers not to crawl URLs under the specified path, subject to the crawler's support for the protocol.
Understanding Allow
The Allow directive can be used in supported implementations to permit crawling of a more specific path inside a broader disallowed area.
Use complex rule combinations only when they are genuinely necessary. Simple configurations are easier to audit and less likely to cause accidental blocking.
Using the Sitemap Directive
A robots.txt file can reference an XML sitemap.
For a sitemap index, the reference might instead point to that sitemap index.
This can provide crawlers with another way to locate sitemap information.
Robots.txt Directives at a Glance
| Directive | Purpose | Example |
|---|---|---|
| User-agent | Identifies the crawler to which a rule group applies | User-agent: * |
| Disallow | Restricts crawling of a specified path | Disallow: /private/ |
| Allow | Permits a more specific path where supported | Allow: /folder/public/ |
| Sitemap | References an XML sitemap location | Sitemap: https://example.com/sitemap.xml |
Robots.txt vs Noindex
Robots.txt and noindex are not interchangeable.
| Method | Main Purpose | Typical Location |
|---|---|---|
| robots.txt | Control crawling | Root robots.txt file |
| noindex | Request exclusion from a search index | Page metadata or supported HTTP response header |
Critical: If you block a page from crawling through robots.txt, the crawler may not be able to access the page and see its page-level noindex directive.
Robots.txt vs Meta Robots
A meta robots directive is placed within an HTML page and can communicate supported indexing and search-result instructions.
By contrast, robots.txt exists outside individual HTML pages and primarily controls crawler access to paths.
Does Disallow Prevent Indexing?
Do not assume that a Disallow rule is a reliable way to remove a URL from search results.
A search engine may know that a URL exists through links or other discovery sources even when crawling of that URL is restricted.
Disallow = Crawl Control
Noindex = Indexing Directive
If the objective is to keep an otherwise accessible page out of search, use an appropriate indexing strategy rather than treating robots.txt as a substitute for noindex.
Can Robots.txt Protect Private Information?
No. Robots.txt is publicly accessible and is not an authentication or security system.
Never rely on robots.txt to secure confidential data. Sensitive resources should be protected with proper authentication, authorization, server configuration, or other appropriate security controls.
What Should You Block in Robots.txt?
There is no universal list of paths every website should block.
Possible candidates on certain websites may include crawlable URL spaces that create large amounts of unnecessary crawler activity without useful search value.
Each rule should have a clear reason. Do not copy another website's robots.txt file blindly.
What Should You Avoid Blocking?
Be especially careful with resources required to understand or render important public pages.
Robots.txt and CSS or JavaScript
Modern webpages may depend on CSS and JavaScript for layout, navigation, content presentation, and rendering.
Blocking resources that search engines need to understand a page can make technical SEO diagnosis more difficult.
Do not block CSS or JavaScript simply because the files themselves do not need to rank. Their purpose may be to help render the pages that do.
Robots.txt Examples
Example 1: Allow General Crawling
An empty Disallow value does not specify a path to block for the matching group.
Example 2: Block a Folder
Example 3: Reference a Sitemap
Example 4: More Specific Access Rule
Always test more complex configurations against the crawler behavior you actually care about.
Robots.txt for WordPress
WordPress sites can expose robots.txt behavior through WordPress itself, plugins, hosting configuration, or a physical robots.txt file depending on the setup.
For WordPress, review:
Robots.txt and Rank Math
If your WordPress website uses Rank Math, robots.txt management should still follow your overall crawling and indexing strategy.
Before changing rules:
Do not add large collections of rules simply because an SEO checklist suggests that robots.txt should look complicated.
Robots.txt for Ecommerce Websites
Ecommerce sites can generate large numbers of URLs through filters, sorting, search, tracking parameters, variants, and navigation systems.
Before blocking these URLs, determine:
Large ecommerce robots.txt changes can affect thousands of URLs at once. Test patterns carefully before deploying them.
Robots.txt for Shopify
Shopify manages parts of its robots.txt behavior through the platform. Customization should be approached carefully because default rules may serve platform-specific purposes.
Before changing Shopify crawling rules, understand which URL pattern you are targeting and why the default behavior is insufficient.
Robots.txt for Staging Websites
Development and staging environments should not rely solely on robots.txt for privacy.
If a staging website must not be publicly accessible, proper access controls are safer than merely asking crawlers not to visit it.
A common migration mistake is carrying restrictive staging settings into the live production website. Always audit crawling and indexing controls after launch.
Common Robots.txt SEO Mistakes
The Dangerous “Block Everything” Mistake
One of the most damaging robots.txt configurations for a public website is an unintended site-wide crawling restriction.
This tells matching crawlers not to crawl the site's paths.
This type of rule may be intentional in a specific controlled environment, but it is usually inappropriate for a public website whose pages you want search engines to crawl.
How to Audit Robots.txt
1. Open the Live File
Check the actual robots.txt file being served by the live website.
2. Identify User-agent Groups
Understand which crawlers each group of rules targets.
3. Review Every Disallow Rule
Determine which URLs match each blocked path and whether the restriction is intentional.
4. Review Allow Rules
Where Allow rules are used, verify that their interaction with broader rules produces the intended crawler access.
5. Check Important Pages
Confirm that service, product, category, blog, resource, and other important public pages are not accidentally blocked.
6. Review Rendering Resources
Check whether essential CSS, JavaScript, images, or other resources have been restricted without a clear reason.
7. Check Sitemap References
If sitemap directives are present, make sure they reference the intended live sitemap locations.
8. Compare With Your Indexing Strategy
Make sure you are not trying to solve an indexing problem solely through a crawling directive.
9. Test After Major Changes
Recheck robots.txt after migrations, staging deployments, CMS changes, plugin changes, or ecommerce architecture changes.
Robots.txt SEO Checklist
Technical Access Is Only the First Step
Once important pages are crawlable and aligned with your indexing strategy, they still need useful content, clear search intent, accurate SEO titles, and relevant meta descriptions.
Try the Free SEO Title & Meta Description GeneratorFrequently Asked Questions About Robots.txt
What is robots.txt in SEO?
Robots.txt is a text file that provides supported web crawlers with instructions about which website paths they may or may not crawl.
Where is the robots.txt file located?
It is normally located at the root of a website host, such as example.com/robots.txt.
Does robots.txt prevent a page from being indexed?
Robots.txt primarily controls crawling and should not be treated as a reliable index-removal method. Indexing controls and crawling controls serve different purposes.
What does Disallow mean in robots.txt?
Disallow specifies a URL path that matching crawlers are instructed not to crawl.
What does User-agent: * mean?
The asterisk acts as a wildcard in a user-agent group and can apply the rules to crawlers that do not have a more specific matching group.
Can robots.txt contain an XML sitemap URL?
Yes. A Sitemap directive can reference the location of an XML sitemap or sitemap index.
Should CSS and JavaScript be blocked in robots.txt?
Not simply because those files do not need to appear as search results. Search engines may need important resources to render and understand public pages properly.
Can robots.txt protect private pages?
No. Robots.txt is publicly accessible and is not a security mechanism. Sensitive resources require proper access controls.
Should every website have a complicated robots.txt file?
No. A simple configuration is often preferable. Add rules only when there is a clear crawling requirement.
Final Thoughts
Robots.txt is a useful technical SEO control, but its role should remain clear: it primarily manages crawler access.
Robots.txt Framework:
Identify Crawler → Identify URL Pattern → Decide Whether Crawling Should Be
Allowed → Add Minimal Rule → Test → Monitor
Avoid using robots.txt to solve unrelated indexing, security, or content quality problems. Keep rules intentional, review them after major website changes, and be especially careful with broad Disallow patterns.
A healthy technical SEO strategy combines appropriate crawling controls with indexing directives, canonicalization, XML sitemaps, internal linking, useful content, and ongoing monitoring.
For page-level optimization, you can also use the Free SEO Title & Meta Description Generator to create starting ideas for your SEO titles and descriptions.