What an llms.txt file does and why you might want one

An llms.txt file is a text document you place in your website's root directory to tell large language models like ChatGPT, Claude, and Gemini how to access and use your content. When an LLM's training process or real-time search encounters your site, it reads this file first to understand what you want indexed, what you want excluded, and how you want your content used.

Think of it as a set of instructions for AI systems, similar to how a robots.txt file works for search engines. The difference is that llms.txt specifically addresses LLM behavior rather than traditional web crawlers. You create it if you want to control which parts of your site LLMs can see, whether they can train on your content, and whether they should cite you when they use your information.

You do not need an llms.txt file for LLMs to find your content — they will discover it anyway through web crawling and training data. But without one, you have no say in how they treat it. With one, you can set boundaries.

Key Takeaways

  • An llms.txt file sits in your website's root directory (the main folder) and contains instructions for how LLMs should handle your content.
  • The file uses a simple text format with sections like User-Agent, Allow, and Disallow to control what LLMs can access and whether they can train on your pages.
  • You can allow certain LLMs while blocking others, require attribution when your content is used, or restrict specific sections of your site.
  • Creating an llms.txt file does not may provide LLMs will follow it, but major models like OpenAI, Anthropic, and Google have stated they respect these files.
  • The file is optional — your site will still be indexed by LLMs without it, but you will have no control over how they use your content.

Where to place the llms.txt file on your website

The llms.txt file must go in your website's root directory, which is the main folder that serves your domain. If your website is example.com, the file should be accessible at example.com/llms.txt. If you run a subdomain like blog.example.com, place a separate llms.txt in that subdomain's root.

For most website platforms, you upload the file the same way you would upload any other file to your public folder. On WordPress, this is usually the root of your hosting account. On Shopify, you may need to use a custom app or contact support, as Shopify does not give direct file access. On static site generators like Hugo or Jekyll, the file goes in your source directory before you build the site.

Once uploaded, test that it is accessible by visiting example.com/llms.txt in your browser. You should see the raw text of the file, not an error page. If you see a 404 error, the file is not in the correct location.

How to write the basic structure of an llms.txt file

An llms.txt file uses a simple text format with sections separated by blank lines. Each section starts with a User-Agent line that names which LLM the rules apply to, followed by directives that tell it what to do.

Here is a basic example that allows all LLMs to access your site:

User-Agent: * Allow: /

The asterisk (*) means "all LLMs." The Allow: / line means "allow access to everything starting from the root." If you want to block all LLMs instead, use Disallow: /.

You can also target specific LLMs by name. Here is an example that allows OpenAI's systems but blocks Anthropic's:

User-Agent: GPTBot Allow: / User-Agent: Claude-Web Disallow: /

Each LLM provider uses a different user-agent name. OpenAI uses GPTBot, Anthropic uses Claude-Web, and Google uses GoogleBot (though Google also respects robots.txt). Check the documentation of the LLM provider to find the exact name they use.

Controlling which pages LLMs can access

You can allow LLMs to see some pages while blocking others by using path-specific rules. For example, if you want to block your admin pages and private content but allow everything else:

User-Agent: * Disallow: /admin/ Disallow: /private/ Disallow: /user-accounts/ Allow: /

The order matters. LLMs read from top to bottom and stop at the first rule that matches. So if you want to allow most of your site but block a few sections, put the specific blocks first, then the general allow at the end.

You can also block specific file types. If you want to prevent LLMs from training on your PDF documents but allow access to your web pages:

User-Agent: * Disallow: /*.pdf Allow: /

The asterisk in *.pdf is a wildcard that matches any filename ending in .pdf.

Setting rules for training and attribution

Beyond blocking access, you can add directives that tell LLMs whether they can train on your content and whether they must cite you. Not all LLMs support these directives yet, but major providers are moving toward respecting them.

Use Disallow-Training: / to prevent LLMs from using your content to train their models, while still allowing them to access it for real-time search or retrieval:

User-Agent: * Allow: / Disallow-Training: /

Use Require-Attribution: true to require that any LLM using your content must cite the source:

User-Agent: * Allow: / Require-Attribution: true

You can combine these directives. For example, allow access and real-time retrieval, but prevent training and require attribution:

User-Agent: * Allow: / Disallow-Training: / Require-Attribution: true

Keep in mind that support for these directives varies. OpenAI and Anthropic have stated they respect training restrictions, but smaller or older LLM systems may not. There is no enforcement mechanism — you are relying on the provider to honor your requests.

Testing and monitoring your llms.txt file

After you create and upload your llms.txt file, verify that it is readable and correctly formatted. Visit example.com/llms.txt in your browser and confirm you see the text you wrote, not an error.

You can also check your website's server logs to see when LLM crawlers access the file. Look for requests from user-agents like GPTBot, Claude-Web, or other LLM identifiers. If you see no requests from LLM crawlers, it may mean they have not discovered your site yet, or they are not respecting the file.

Some LLM providers publish transparency reports or documentation about which sites they crawl. OpenAI publishes a list of IP addresses used by GPTBot, and you can cross-reference your logs against that list. Anthropic and Google provide similar information in their documentation.

If you want to monitor whether LLMs are following your rules, you can add a tracking pixel or unique URL to a section you have blocked, then check whether that URL appears in LLM-generated content. If it does, the LLM is not respecting your directives. Report this to the provider.

What happens if you do not create an llms.txt file

Without an llms.txt file, LLMs will still find and index your content through normal web crawling and training data collection. They will treat your site as if you have allowed everything — all pages are accessible, all content can be used for training, and no attribution is required.

This is the default behavior and is how most websites are currently handled. LLMs do not wait for permission; they assume access unless told otherwise. If you want any control over how your content is used, creating an llms.txt file is the current standard way to express that preference.

However, creating a file does not may provide compliance. Smaller LLM systems, older models, or systems trained on older data may not check for or respect llms.txt files. The file is a signal of your intent, not a legal requirement or technical barrier.

Frequently Asked Questions

Is llms.txt the same as robots.txt?

No. robots.txt controls traditional search engine crawlers like Googlebot and Bingbot. llms.txt is specifically for large language models. You can have both files on your site, and they can have different rules. For example, you might allow Google to index your site but block LLMs from training on your content.

Do all LLMs respect llms.txt files?

Major providers like OpenAI, Anthropic, and Google have stated they respect llms.txt files, but smaller or independent LLM systems may not. There is no enforcement mechanism, so you are relying on each provider to honor your directives. If you discover an LLM ignoring your file, contact the provider directly.

Can I use llms.txt to prevent my content from appearing in LLM responses?

Partially. You can block LLMs from training on your content with Disallow-Training, which prevents it from being used in future model updates. However, LLMs trained before your file was created will still have your content in their weights. You cannot retroactively remove content from already-trained models.

What if my website is on a subdomain or hosted on a platform that does not allow file uploads?

Place an llms.txt file in the root of each subdomain where you want rules to apply. If your platform does not allow direct file uploads, contact support to ask if they can add the file for you, or consider using a different hosting provider that gives you file access.

Should I block all LLMs or allow them?

That depends on your goals. If you want your content to appear in LLM responses and reach users that way, allow access. If you want to protect proprietary or sensitive content, or if you want to require attribution, set specific rules. There is no single right answer — it is your choice based on your content and business model.