llms.txt
By Olha Kryva ·
llms.txt is a plain-text file describing what your site is and pointing to its main pages in Markdown. It is a proposal, not a standard, and no major assistant has confirmed reading it. It costs an hour, it is harmless, and it is not a ranking factor — good housekeeping rather than a lever.
On this page 16 sections
- The short version
- What is llms.txt?
- What it is not
- Why write one anyway
- How to write one worth having
- What goes in llms-full.txt
- Does it actually get read?
- Where it sits in the priority list
- A worked example
- Frequently asked
- What to put in the file, line by line
- How to keep it from going stale
- Should you block AI crawlers instead?
- Questions people actually ask
- Where to start
- One thing to do tomorrow
The short version
llms.txt is a plain-text file at the root of a site that describes what the site is and points to its main pages in Markdown. It is a proposal, not a standard: no major assistant has publicly committed to reading it.
It costs an hour to write, it is harmless, and it is not a ranking factor. Treat it as good housekeeping rather than a lever, and be sceptical of anyone selling it as the latter.
What is llms.txt?
A file at /llms.txt containing:
- a short description of what the site is and who it is for;
- its key pages, each with a one-line summary;
- optionally a link to
/llms-full.txt— the full text of those pages in Markdown, without navigation, forms or interface furniture.
The idea is that a model retrieving your site gets a clean, unambiguous summary instead of parsing a page full of menus, cookie banners and calls to action.
It was proposed in 2024 and adopted fairly widely by developer-facing sites and documentation projects, where the use case is strongest: a model asked how to use a library benefits enormously from clean Markdown docs.
What it is not
Not a standard. It is a proposal with community adoption. Whether the large assistants read it is, as of now, not publicly confirmed by any of them.
Not robots.txt. It grants nothing and blocks nothing. Crawling permissions still live in robots.txt, and llms.txt has no effect on them whatsoever. This confusion is common and worth stating plainly: publishing llms.txt does not allow or disallow anything.
Not a ranking factor. Nobody has demonstrated a measurable visibility change from adding one — us included, and we measure this for a living. Anyone who tells you otherwise has a file to sell you.
Not a sitemap. A sitemap tells a crawler which URLs exist. llms.txt tells a reader what the site is about. Different jobs; you want both.
Why write one anyway
Three honest reasons, none of them dramatic.
1. It forces you to state plainly what you do. Writing two sentences that a stranger would understand, then listing your pages with one line each, is an exercise most sites fail the first time. Whatever the file’s fate, that exercise improves the site.
2. It is a place to find contradictions. If your homepage and your pricing page describe your offer differently, writing this file makes it obvious in a way nothing else does. Contradiction between sources is a genuine reason for a model not to trust any of them, and it is the most common problem we find in audits.
3. The cost is an hour. For a category this young, cheap and harmless is a reasonable bet. We would not spend a day on it, and we would not pay anyone to write one.
How to write one worth having
Describe the site in two or three sentences a stranger would understand. No adjectives if you can avoid them; say what it does and for whom.
List the pages that matter, each with a one-line summary. Not every URL — the ten or fifteen pages someone would actually need. A file listing four hundred URLs has become a sitemap with worse formatting.
Keep it factually identical to the site. A file saying something the site contradicts is worse than no file. This is the rule that matters most, and it is the one people break by writing the llms.txt six months before the last site redesign.
Include dates where content is dated. A retrieval system reading it benefits from knowing what is current.
Leave out invented examples and demo data. If your homepage has made-up company names for illustration, keep them out of the file — an agent may read them as real businesses. We found exactly this on our own site during an audit, and it is a subtle, embarrassing failure mode.
Link /llms-full.txt from it if you publish one, in an “Optional” section.
Update it when pages change. A stale file is a contradiction generator.
What goes in llms-full.txt
The full text of your main pages in Markdown, stripped of everything that is interface rather than content: navigation, forms, cookie notices, repeated calls to action, decorative lists.
Two rules that matter:
Generate it from the pages themselves, not by hand. A hand-written version drifts from the site within one redesign, and then you have published a contradiction.
Strip demo content. See above.
Ours is at /llms-full.txt, generated from the same page content the site renders, which is the only way we have found to keep the two in agreement.
Does it actually get read?
The honest answer is that we do not know, and neither does anyone claiming otherwise.
What can be observed. Server logs show which user agents request the file. If you publish one, check whether anything fetches it — that is the only direct evidence available to you, and it is worth looking at before believing a vendor.
What cannot be observed. Whether a model that fetched it weighed it, or whether an answer naming your business was influenced by it. There is no attribution path from an assistant’s answer back to a file it may have read.
What we can say from measurement. We have not been able to show a visibility change attributable to publishing one. That is not proof it does nothing; it is a statement about the size of any effect relative to the noise in this data, which is several points.
Where it sits in the priority list
Below almost everything.
| Work | Effect on being named | Effort |
|---|---|---|
| Making your own pages answer buying questions directly | High | Days |
| Correcting how you are described in cited sources | High | Weeks, other people’s queues |
| Publishing something checkable | Medium, durable | Weeks |
| Schema markup | Low, helps parsing | Hours |
| llms.txt | Unproven | One hour |
Write it on a slow afternoon. Then go and work on the pages that actually decide the answer: your own service pages first, because they are the only description of you that you write, and then the directories and review sites read alongside them. In one market we measured, eleven pages stood behind an average answer, and they have to agree with each other.
A worked example
From our own site, because it is the one we can show in full.
What we published. An llms.txt with a description of the service, the method in four lines, the pages that matter with one-line summaries, a short Q&A, contact details with working hours, and a link to llms-full.txt.
What the audit found afterwards. Two contradictions we had not noticed. The file described the sample report as belonging to a fictional company while the homepage described it as a real client’s report published with permission. And the file promised a reply within one business day while the contact page promised ten minutes in working hours.
Which is the point. Neither contradiction was caused by the file; both were already on the site and had been for weeks. Writing the file is what surfaced them. We fixed both on the site and in the file, and the working hours are now stated with a timezone.
What it did for visibility. Nothing we can measure.
Frequently asked
Do I need llms.txt to appear in AI answers? No. There is no evidence it affects whether you are named.
Will it hurt me? No, unless it contradicts your site.
Should I block AI crawlers instead? Different question entirely, and it belongs in robots.txt. Blocking removes you from retrieval-based answers, which for a business wanting to be recommended is usually the wrong direction.
Who reads it today? Some developer tooling and documentation systems demonstrably do. Whether the large consumer assistants do is unconfirmed.
Should I pay someone to write one? No. It is an hour, and the value is in the thinking, which you cannot outsource without losing the point of it.
What to put in the file, line by line
A worked structure, in the order that makes sense to something reading it once.
1. The name and one sentence. What the site is, for whom. No adjectives.
2. What you do, in three or four lines. Concretely: what the service measures, covers, or sells. This is the part an assistant may quote.
3. How you are careful. If your method has limits — ours does — state them here. A source that names its own error bars is easier to trust than one that does not, and this is the cheapest place to do it.
4. The pages, with one line each. Ten to fifteen. Each line should tell a reader whether to open the page.
5. A short question-and-answer section. The five questions people actually ask before buying. Same wording as on the site.
6. Contact, with working hours and a timezone. “We reply quickly” is not information. “A person replies within ten minutes, 08:00–20:00 Kyiv time” is.
7. An optional section linking /llms-full.txt and anything else
machine-readable.
How to keep it from going stale
The failure mode is not writing a bad file; it is writing a good one and letting the site move away from it.
Generate llms-full.txt from the pages. Not by hand. Ours is built from the same page content the site renders, which is the only way we have found to keep the two in agreement through a redesign.
Diarise a quarterly check of the short file against the site. Ten minutes.
Update it when you change a promise. Response times, prices, what is included. These are exactly the facts a model will cross-check against your pages, and a contradiction costs more than the file gains.
Treat it as published copy, not configuration. Whoever approves site copy should approve this.
Should you block AI crawlers instead?
A different question, often asked in the same breath, and worth separating.
Blocking a crawler in robots.txt removes your pages from that system’s retrieval. For a publisher whose business is paid page views, that can be a rational decision. For a business that wants to be recommended, it is close to opting out of the recommendation.
The asymmetry worth understanding: blocking protects your content from being summarised, and it also prevents your business from being described accurately when someone asks about your category. Most service businesses want the second more than they fear the first.
If you do block, be deliberate about which agents and check what it costs you by measuring before and after. Most people who block do it by copying someone else’s robots.txt, which is how sites end up blocking crawlers they meant to welcome.
Questions people actually ask
Is llms.txt the same as robots.txt? No. robots.txt controls crawling; llms.txt describes the site. Publishing one does not allow or disallow anything.
Does anyone check whether I have one? Some developer tooling does, demonstrably. Whether the large consumer assistants do is unconfirmed by any of them.
Should it list every page? No. Ten to fifteen pages with one line each. A file listing four hundred URLs is a sitemap with worse formatting.
How long should llms-full.txt be? As long as your main pages are, stripped of navigation and forms. Generate it rather than writing it.
Where to start
Write one. It will take an hour and it will find you at least one contradiction on your own site.
Then spend the rest of the week on something that moves the number: how to improve AI visibility puts the work in order, schema markup covers the other machine-readable file people ask about, and the free check tells you where you actually stand before you spend anything.
One thing to do tomorrow
Write the two sentences that describe your site to a stranger. If that takes more than ten minutes, the file is not your problem — the positioning is, and no machine-readable format will fix it. That exercise is most of the value here, and it is available whether or not you publish anything.
Questions
- Is llms.txt an official standard?
- No. It was proposed in 2024 and adopted by a number of developer-facing sites. Whether the large assistants read it is not publicly confirmed by them.
- Does it replace robots.txt?
- No. It grants nothing and blocks nothing. Crawling permissions still live in robots.txt.
- Will it improve my AI visibility?
- Nobody has shown a measurable change from adding one, us included. Write it because stating plainly what you do improves the site either way.