# llms.txt

llms.txt is a Markdown file at /llms.txt that tells language models what your site is, what it holds and what can be skipped. You use it to curate links and summaries for agents, not to block crawlers or set licences. For a small site, it is an hour’s job that can reduce irrelevant fetches and steer better citations.

Updated 2026-09-13 · Source: https://porteur.ai/glossary/llms-txt

## What llms.txt is and why it matters

llms.txt is a proposal from September 2024 by Jeremy Howard of Answer.AI at llmstxt.org. It is a plain Markdown file served at /llms.txt. It gives AI models and agents a curated map of your site.

It has one required element: an H1 with your site or project name. Then a blockquote with a one paragraph summary. Optional paragraphs may follow. After that, H2 sections each hold a list of links in the form [title](url): one line description. A section titled Optional marks links an agent can skip when its context is short. /llms-full.txt can optionally carry the full text of pages as Markdown.

It is not access control, an opt out or a licence file. robots.txt governs what a crawler may fetch. llms.txt says what the site is and where the useful pages are.

> Adoption is voluntary. As of 2026, no major AI vendor has committed to reading it.

## The format by example

```md
# YourProduct

> A tool for small teams to track bugs and ship fixes. Web app and API. Free tier. Docs, changelog and pricing below.

We prioritise stability and clear docs.

## Docs
- [Getting started](/docs/getting-started): Install, invite your team, first project in 10 minutes.
- [API reference](/docs/api): REST endpoints with examples.
- [Webhooks](/docs/webhooks): Events you can subscribe to.

## Product
- [Pricing](/pricing): Plans and limits.
- [Changelog](/changelog): What shipped this month.
- [Status](https://status.yourproduct.com): Uptime and incidents.

## Company
- [About](/about): Team and contact.
- [Security](/security): Data handling and audits.

## Optional
- [Blog](/blog): Commentary and tutorials.
```

A good file reads like a tidy menu, not a dump. Keep link titles human. Keep descriptions to one clean line each. Put long lists in their own H2.

## How to create and publish your llms.txt

1. **Draft the outline** Write the H1 as your site name: "YourProduct". Add a one paragraph summary in a blockquote. Add one or two plain paragraphs if they help.
2. **Group key pages** Create H2 sections like Docs, Product, Company. List 5 to 15 links that define the site. Format each as [Title](URL): one line description.
3. **Mark optional links** Add an H2 named Optional for nice to have links, like blog archives or long indexes.
4. **Save and serve it** Save as UTF 8 without BOM. Name it llms.txt. Place it at the web root so it resolves at https://yourproduct.com/llms.txt.
5. **Consider /llms-full.txt** If you want agents to ingest full content, publish a separate /llms-full.txt with Markdown copies of key pages.

On a static site, add llms.txt to the public folder. On a framework, map a static route to /llms.txt with text/plain or text/markdown. Avoid redirects.

## How to test and keep it useful

- Open /llms.txt in a browser. Check the H1, blockquote, H2s and the link lines render as plain text.
- Click a few links. They must be absolute or correct relative URLs. No 404s. No redirects if you can avoid them.
- View the response headers. text/plain or text/markdown is fine. Status 200. Cache it for a day or a week.
- Validate by eye against llmstxt.org. Confirm the Optional section is only things an agent can skip.
- Set a quarterly reminder. Update when you add a major page like /pricing or /docs/api.

For /llms-full.txt, spot check the Markdown renders, headings are clear and code blocks are fenced. Keep it in sync with live pages after big edits.

## Traps to avoid

- Treating llms.txt as a crawler gate. It is not. Use robots.txt for access control.
- Listing everything. Curate. Ten strong links beat fifty weak ones.
- Forgetting the blockquote summary. Many agents will read that first.
- Letting URLs rot. A single stale /docs link penalises trust by agents and users.
- Burying critical info under Optional. Keep docs, pricing and status in main sections.
- Serving HTML. Keep it plain Markdown at /llms.txt, not a webpage.

## Questions

### What is an llms.txt file?

It is a Markdown file at /llms.txt that gives AI models and agents a curated map of your site. It has an H1 with the site name, a one paragraph blockquote summary, optional paragraphs, then H2 sections with lists of [title](url): one line description links. It is not an access control or licence file.

### Is LLMs.txt actually used?

Adoption is voluntary. As of 2026, no major AI vendor has committed to reading it. You can still publish it now, as it costs about an hour and may help early agent tooling.

### Is llms.txt mandatory?

No. It is a proposal, not a standard you must follow. If you skip it, crawlers can still discover your pages through links and sitemaps. You publish it to improve how agents understand your site.

### How do I generate an llms.txt file?

You can write it by hand in any text editor following the format on llmstxt.org. Keep the H1, blockquote and H2 list sections. If you use a generator, review every link and description before publishing.

### Where should I put llms.txt on my site?

Serve it at the root so it is reachable at https://yourdomain.com/llms.txt. Avoid redirects. Set a cache that suits your update cadence, for example a day or a week.

### How is llms.txt different from robots.txt?

robots.txt tells crawlers what they may fetch. llms.txt tells models what your site is and which pages matter. Use both: robots.txt for access, llms.txt for curation.

## Read next

- [llms.txt generator](https://porteur.ai/tools/llms-txt-generator): Write an llms.txt from your site or from a few fields, in sections, and check the one a site already has against the proposal, link by link.
- [robots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot and what to allow](https://porteur.ai/guides/robots-txt-for-ai-crawlers): Decide which AI crawlers to allow in robots.txt, why it matters, and copy‑paste examples for GPTBot, ClaudeBot, PerplexityBot, Google‑Extended and more.
- [AI crawler](https://porteur.ai/glossary/ai-crawler): An AI crawler is a bot that fetches your pages for an AI. See the main bots, how to spot them in logs, and how to allow or block them well.
- [Generative engine optimization (GEO): how to be named by AI answers](https://porteur.ai/guides/generative-engine-optimization): How to get your site named in AI answers: sources, crawlers, lists, structure, llms.txt, comparisons, and how to measure by asking assistants.
- [How to appear in Google’s AI Overviews](https://porteur.ai/guides/how-to-rank-in-ai-overviews): What overviews cite, which queries trigger them, and how to structure pages that get linked. Steps to find your openings and measure them.
- [AI crawler access checker](https://porteur.ai/tools/ai-crawler-access-checker): Paste a site or a page and see, crawler by crawler, whether it may read it: root and page, the rule that decides, the text without JavaScript, llms.txt.

Get a quick read on where your /llms.txt would help most: Porteur scans your URL, the searches around it and rival pages in about thirty seconds and shows three findings free. Free check: https://porteur.ai/
