llms.txt vs robots.txt: a reading list and a crawl policy
robots.txt (RFC 9309) tells crawlers which paths they may fetch. llms.txt, a community proposal, points models to a site's best content. How they differ.
robots.txt is for telling automated crawlers which parts of a site they may fetch. llms.txt is for giving a language model, usually one helping a user at the moment of a question, a short, curated Markdown guide to a site’s most useful content.
One file is a policy about access and the other is a reading list. They sit side by side at the root of many sites, and neither does the other’s job.
Status as of September 26, 2026. robots.txt is an IETF standard: RFC 9309, the Robots Exclusion Protocol, published in September 2022 on the Standards Track. llms.txt is a community proposal. Jeremy Howard published it at llmstxt.org on September 3, 2024, and the site now presents a revised v2, dated August 2026. A search of the IETF Datatracker finds no document for it, and no standards body has adopted it. Separately, the IETF AI Preferences working group, chartered in January 2025, is drafting a way to attach AI usage preferences to robots.txt; its drafts are still Internet-Drafts.
What robots.txt is for
RFC 9309 standardizes the method Martijn Koster first defined in 1994. Service owners use it to keep crawlers out of parts of their URI space. The essentials:
- Location. The file is
/robots.txt, all lowercase, in the top-level path of the service, UTF-8 encoded and served astext/plain. Each scheme and authority has its own file. - Groups. Rules are grouped under
user-agentlines. A crawler matches its product token case-insensitively and obeys the matching group, or the*group if none matches. - Rules.
allowanddisallowtake path patterns. The most specific match, the one with the most octets, wins. If an allow and a disallow rule are equivalent, the allow rule should be used. - Errors. A 4xx response means the file is unavailable, and the crawler may access anything. A 5xx or network error means it is unreachable, and the crawler must assume everything is disallowed. After a long outage, 30 days in the RFC’s example, the crawler may treat the file as unavailable or keep using a cached copy.
- Limits. Crawlers should not use a cached copy for more than 24 hours unless the file is unreachable, and must parse at least 500 KiB.
Illustrative:
User-agent: ExampleBot
Disallow: /checkout/
Allow: /checkout/help
User-agent: *
Disallow: /account/
The RFC is explicit about its limits. The rules are ones crawlers are requested to honor, and they are not a form of access authorization. Its security section adds that listing paths in robots.txt makes them public, so real protection needs real access controls.
What llms.txt is for
The llms.txt proposal starts from a practical problem. Web pages wrap their content in navigation, scripts and layout, and a whole site will not fit in a model’s context. A small Markdown file at a predictable location lets an agent that is helping a user find the right material quickly. The proposal expects it to be most useful at inference time, when an agent looks something up, rather than for training.
The v2 format:
- The file sits at
/llms.txtor any subpath, such as/docs/llms.txt, and covers the pages under that path. Where several apply, the most specific wins. - An H1 with the project or site name is the only required section.
- A blockquote gives a short summary, followed by optional detail.
- H2 sections hold lists of Markdown links with optional notes.
- A section titled “Optional” marks links an agent can skip when its context is short.
v2 also proposes clean Markdown copies of pages at the page URL with .md added or the extension replaced, and link relations so agents can find them. Illustrative:
# Example Freight
> Example Freight quotes and books less-than-truckload shipments between Canada and the US.
## Docs
- [Shipping guide](https://example.com/docs/shipping.md): Pallet sizes, lanes and pickup windows
- [Rates FAQ](https://example.com/docs/rates.md): How quotes are calculated
## Optional
- [Company history](https://example.com/about.md)
The proposal itself contrasts the two files. robots.txt tells automated tools what access is acceptable. llms.txt supplies information on demand. It says nothing about permission to crawl or to train.
Side by side
| robots.txt | llms.txt | |
|---|---|---|
| Purpose | Tell crawlers which paths they may fetch | Point models to a site’s most useful content |
| Layer | Crawl policy for automated clients | Content guidance for model context |
| Who talks to whom | A crawler reads the site owner’s rules before crawling | A model or agent reads the owner’s curated list while helping a user |
| Transport | HTTP GET of /robots.txt, text/plain |
HTTP GET of /llms.txt or a subpath; Markdown |
| Discovery | Fixed location at the top-level path of each host | Fixed filename at the root or any subpath; link relations in v2 |
| Auth | None; the RFC says it is not access authorization | None; a public reading list |
| State | Crawlers cache it, normally for up to 24 hours | None defined |
| Governance and status | IETF RFC 9309, Standards Track (September 2022) | Community proposal (llmstxt.org), v2 dated August 2026 |
Where AI usage preferences fit
Neither file, as published, lets a site say “you may index this for search but not train models on it.” The IETF AI Preferences working group is drafting that. Its vocabulary draft (draft-ietf-aipref-vocab-08, September 2026) defines usage categories with short labels: train-ai for AI training, ai-use for using content as input to a generative model, and search. Its attachment draft (draft-ietf-aipref-attach-05, August 2026) adds a Content-Usage rule to robots.txt groups and an HTTP header of the same name, for example Content-Usage: train-ai=n. The attachment draft would update RFC 9309 if approved. Both are works in progress and can change before publication.
So the likely home for AI usage preferences is robots.txt and HTTP headers. llms.txt stays a reading list.
When to use each
Publish robots.txt on every site that crawlers reach. Use it to keep crawlers away from paths with no value to them, such as carts, internal search results and account pages, and to address specific crawlers by product token. Do not use it to hide anything sensitive.
Publish llms.txt when you want agents helping your users to find your best documentation, product pages or policies quickly, and when you can keep the list current. It is cheap to try. Treat it as a proposal: some agents will read it and others will not.
Using both
They work together if they agree. If llms.txt links to Markdown pages under a path that robots.txt disallows for the crawlers you care about, those crawlers will skip them. Check the two files against each other whenever either changes.
If you also run an A2A agent, llms.txt can link to your Agent Card so a model reading your docs learns that a structured endpoint exists. Agent Card vs llms.txt covers that pairing.
Common misconceptions
“robots.txt blocks bots.” It states rules that well-behaved crawlers follow. The RFC says it is not access authorization. Rate limits, authentication or signed-request checks such as Web Bot Auth do the enforcing.
“llms.txt is a standard like robots.txt.” It is a community proposal with no standards-body adoption. It borrows the fixed-filename idea from robots.txt and describes itself as serving a different purpose.
“Putting a page in llms.txt grants permission to train on it.” The proposal makes no statement about permissions. Usage preferences belong to the AI Preferences drafts.
“robots.txt covers every automated request.” RFC 9309 addresses crawlers, which it describes as automated clients that traverse links. It does not say how an agent that fetches one page at a user’s request should treat the file, so operators document their own practice.
Questions
- Does llms.txt say whether AI companies may train on my content?
- No. The proposal describes llms.txt as information an agent uses on demand while helping a user, and it makes no statement about permission to crawl or train. Usage preferences for AI are the subject of the IETF AI Preferences working group, whose drafts extend robots.txt and HTTP headers.
- Can llms.txt live somewhere other than the site root?
- Yes. The v2 proposal allows /llms.txt or any subpath, such as /docs/llms.txt, and the most specific file covers the pages under it. robots.txt must sit at the top-level path of the host, and its rules apply to that scheme and authority only.
Sources
- RFC 9309: Robots Exclusion Protocol (September 2022) (accessed )
- The /llms.txt file, v2 (Jeremy Howard; first published September 3, 2024) (accessed )
- llms.txt: Changes, v2 (August 2026) (accessed )
- IETF AI Preferences (aipref) working group (accessed )
- draft-ietf-aipref-attach-05: Associating AI Usage Preferences with Content in HTTP (August 19, 2026) (accessed )
- draft-ietf-aipref-vocab-08: A Vocabulary For Expressing AI Usage Preferences (September 14, 2026) (accessed )
- IETF Datatracker document search for llms.txt (no matching documents) (accessed )