What the file contains
A short description of the site, then a curated list of the URLs worth reading, usually grouped with a line of context each. The point is to hand a model a clean map instead of making it infer structure from navigation and internal links.
Some sites also publish llms-full.txt, which inlines the actual content of those pages as markdown so a model can read the substance without fetching anything.
| robots.txt | sitemap.xml | llms.txt | |
|---|---|---|---|
| Status | Long-standing standard | Long-standing standard | Proposed convention |
| Used by Google | Yes | Yes | No |
| Controls | Crawler access | Discovery of URLs | Nothing, it advises |
| Worth publishing | Essential | Essential | Cheap, unproven |
Be honest about what it does
Adoption is genuinely unclear. No major search engine has confirmed it as a ranking or citation input, and Google has said publicly that Search does not use it. Treating it as a lever that lifts AI visibility on its own is not supported by evidence.
It is cheap to publish and does no harm, which is a reasonable argument for doing it. It is not a reason to bill it as a deliverable that moves citations.
What actually drives citation
- Letting the AI crawlers reach your pages at all, which robots.txt does control
- Server-rendered HTML, since content behind client-side JavaScript may never be read
- Passage structure that survives extraction
- Third-party mentions, which weigh more than anything on your own domain
Why it matters for B2B marketing teams
We mention this one mainly to save you money. If an agency proposes llms.txt as a deliverable that will lift your AI visibility, ask for the evidence. It takes ten minutes to publish, which makes it a fine housekeeping task and a poor line item.

