What Is llms.txt? Does Your Website Need One in 2026?

You ship a new feature, update your docs, and then watch an AI assistant answer questions about your product using an old help article you buried three clicks deep. llms.txt is an attempt to fix that problem with a simple idea: publish a short, curated “read this first” list for language models.

llms.txt is a Markdown file (usually at /llms.txt) described in the llms-txt v2 proposal. It points AI agents to the pages you consider your source of truth—documentation, product explanations, policies, pricing pages, and similar. Keep your expectations grounded. Ahrefs notes that no major LLM provider currently supports llms.txt, including OpenAI, Anthropic, and Google (source), so this is not something you should treat as a ranking or citation tactic.

Google’s Lighthouse team calls llms.txt an “emerging convention” for agentic browsing, and it treats the file as optional. If /llms.txt returns a 404, Lighthouse marks the audit as Not Applicable (N/A) (Chrome for Developers). That’s the right mental model: llms.txt is optional housekeeping that can help agents choose the right page faster, but it can’t override robots.txt, make blocked pages crawlable, or rescue thin content.

This guide shows what llms.txt looks like in practice, when it’s worth adding (and when it’s busywork), and what to check first so your “recommended” URLs aren’t pointing to locked doors.

What Is llms.txt?

An llms.txt file is a simple way to publish that “curated reading list” in a format language models can scan quickly. The llms-txt v2 proposal describes it as a Markdown file named llms.txt, commonly hosted at /llms.txt, that points LLMs and AI agents to the pages you consider most useful.

In plain terms: llms.txt is a human-maintained index of your best content for machine readers. You pick the URLs, group them into sections, and add short notes so an agent can choose the right page without guessing.

The proposal focuses on two practical goals:

  • Reduce ambiguity for AI agents that land on busy navigation pages, tag archives, or search results.
  • Prioritize canonical, high-signal pages such as product docs, API references, pricing, policies, and “getting started” guides.

llms.txt does not replace normal SEO plumbing. It does not grant permission to crawl URLs, and it does not force indexing. It simply lists what you want an agent to read first, assuming the agent can fetch those pages.

Where llms.txt Lives and What It Covers

The llms-txt v2 proposal says you can place the file at the site root (/llms.txt) or at a subpath such as /docs/llms.txt, and it “covers” the pages under that path. For example, /docs/llms.txt describes everything in /docs/ (llms-txt v2).

This scoping matters on large sites. You can publish separate files for separate areas, and the proposal says that when more than one llms.txt applies, agents should use the most specific one (llms-txt v2).

If you want a second opinion on whether this is worth doing, keep the definition simple: llms.txt is a Markdown map of your best pages for LLM consumption, maintained by you, published at /llms.txt or a relevant subpath, and governed by the llms-txt v2 proposal.

llms.txt Format Rules (With a Copy-Paste Example)

The llms.txt format stays simple on purpose. The llms-txt v2 proposal treats it as a Markdown reading list for agents and LLMs, so you can write it in any editor and serve it as plain text from /llms.txt.

Per the spec, only one element is required: an H1 that names your site or project. Everything else is optional, but a consistent structure makes the file easier for a machine to skim.

  • Required: An H1 (a single line starting with #) with your site or project name. The spec also allows an optional byte-order mark (BOM).
  • Optional: A short blockquote summary (lines starting with >) that gives the context needed to interpret the links.
  • Optional but recommended: One or more H2 sections (lines starting with ##) that group links into “file lists.”
  • Link pattern: Each bullet in a file list uses a required Markdown link, then optionally : plus notes about what the page contains.

llms.txt Example You Can Copy and Edit

This llms.txt example follows the spec’s shape: H1, optional summary, then H2-delimited link lists with short notes. Replace the URLs and labels with your real “source of truth” pages.

# ExampleCo

> Official documentation and policies for ExampleCo. Use these pages as the primary sources.

Documentation

Policies

Keep notes short and factual. If you cannot describe the page in one sentence, the page itself probably needs a clearer structure before you add it to llms.txt.

Free tool: the llms.txt generator creates a file in the llmstxt.org format from your key pages.

How Is llms.txt Different From robots.txt and sitemap.xml?

That “one-sentence note” mindset is a good way to spot what llms.txt is and is not. llms.txt is a curated list of pages you want an LLM to read first. It does not control crawling, and it does not help search engines discover everything on your site. That job belongs to robots.txt and sitemap.xml.

Here’s the clean mental model:

  • robots.txt answers: “May a crawler fetch this URL?”
  • sitemap.xml answers: “Here are the URLs that exist and matter for discovery.”
  • llms.txt answers: “Start with these pages, they explain the site best.”

robots.txt: Permissions and Crawl Boundaries

robots.txt controls access. If you disallow a path in robots.txt, an agent that respects robots.txt should not fetch it, even if you list it in llms.txt. In practice, that means llms.txt cannot “fix” a blocked docs area, a gated pricing page, or an API reference hidden behind restrictive rules.

Use robots.txt for things like blocking admin paths, internal search results, staging environments, or parameter-heavy URLs. Use llms.txt for pointing to your canonical “read this first” pages, assuming they are crawlable.

Free tool: the robots.txt generator builds a file with your sitemap and blocked paths, plus presets for blocking AI training bots while keeping AI search.

sitemap.xml: Discovery and Indexing Hints

sitemap.xml helps search engines find URLs. You typically use it to list canonical pages you want indexed, at scale. Google Search Console and Bing Webmaster Tools can consume sitemaps directly, and many CMS platforms can generate them automatically (WordPress plugins like Yoast SEO or Rank Math, or built-in sitemaps in modern setups).

llms.txt is different in intent. It is selective. A sitemap might list thousands of URLs, while llms.txt usually lists the handful that explain your product, docs, policies, and terminology clearly.

llms.txt: Curation Without Enforcement

llms.txt curates, it does not grant access or guarantee use. The llms-txt v2 proposal describes a Markdown file at /llms.txt (or a subpath) that points models to high-signal pages. It does not override robots.txt, and it does not function like a sitemap for indexing.

A quick self-check: if you add a URL to llms.txt, verify two things. Your server returns 200 OK for that URL, and robots.txt allows crawling for the relevant user agents. If either fails, the recommendation in your llms.txt becomes dead weight.

Do I Need llms.txt? Use This Quick Decision Checklist

If your recommended URLs return 200 OK and robots.txt allows crawling, the remaining question is simple: do you need llms.txt at all? For many sites, llms.txt is optional housekeeping. It makes the most sense when you have a lot of “right answers” spread across many pages.

Use this checklist as a yes or no filter. If you hit several “yes” items, adding llms.txt is usually worth the small maintenance cost. If you hit several “no” items, skip it and focus on the basics covered later.

  • Yes: You publish documentation, API references, help center articles, or a large knowledge base where an agent can easily land on the wrong page.
  • Yes: Your site has multiple versions of the same answer (blog post, docs page, release notes), and you want to point to a single canonical source.
  • Yes: You ship frequent changes (features, pricing, policies) and you already maintain a changelog, release notes, or a docs index.
  • Yes: You can name 10 to 30 “source of truth” URLs today (getting started, core concepts, API auth, pricing, security, privacy policy) without debating it for a week.
  • Yes: Someone on your team can own link hygiene, meaning they will update llms.txt when URLs change or content moves.
  • No: Your site is mostly a brochure, a portfolio, or a small marketing site with a handful of pages. Navigation already tells the whole story.
  • No: Your best content sits behind a login, blocks bots, or lives in PDFs you do not want crawled. llms.txt cannot grant access.
  • No: You cannot keep it current. A stale llms.txt sends agents to outdated docs and retired features.

When llms.txt Is Most Likely to Help

llms.txt works best as a curated map for documentation-heavy sites. It gives an AI agent a fast path to the pages you would send a human to, instead of forcing the agent to guess from navigation, on-site search, or tag archives.

Keep expectations realistic. Even in 2026, llms.txt remains a proposal, and Ahrefs notes that no major LLM provider currently supports llms.txt (including OpenAI, Anthropic, and Google). Treat it as a low-risk hint, not a ranking tactic.

How to Create an llms.txt in 10 Minutes

If you decide to publish llms.txt, treat it like a maintained reading list. The goal is simple: put your highest-signal pages in one Markdown file so an agent can pick the right source fast. This takes minutes if you keep the scope tight and only link to pages that already return 200 OK.

  1. Pick 5 to 15 “source of truth” URLs. Favor pages that explain your product or topic without requiring clicks through navigation. Good candidates include: “Getting started,” API reference, core feature docs, pricing, security, privacy policy, and a glossary. Skip tag archives, thin blog posts, and pages that change daily.

  2. Confirm every chosen URL is fetchable. Open each page in an incognito window. Make sure it loads without login, geo-blocking, or a cookie wall. If you block the path in robots.txt, remove it from llms.txt or fix the block first.

  3. Draft the file in plain Markdown. Follow the llms-txt v2 proposal: start with one H1 (required), optionally add a short blockquote summary, then group links under H2 sections. Keep notes short and factual.

  4. Name it exactly llms.txt. Use UTF-8. Avoid smart quotes and odd formatting from rich text editors.

  5. Publish at the right path. Put it at /llms.txt for site-wide coverage, or at a subpath like /docs/llms.txt if you only want it to describe that area. The spec allows both (llms-txt v2).

  6. Serve it as a normal, public file. Your server or CDN should return 200 OK. Do not redirect to HTML. Do not require authentication. On WordPress, you can add a static file at the web root via your host’s file manager, SFTP, or a deployment pipeline. On Webflow, you can typically upload it as a static asset and route it to /llms.txt using redirects or hosting configuration, depending on your setup.

  7. Verify it returns 200 OK and the content matches what you wrote. Load /llms.txt in your browser, then check the response in DevTools (Network tab) or with curl -I https://yoursite.com/llms.txt. Fix 404, 500, or unexpected redirects.

If you want a quick sanity check, Balzac’s free AI crawler checker can confirm whether /llms.txt exists and whether common AI crawlers can reach your site: https://hirebalzac.ai/free-seo-tools/ai-crawler-checker/.

What To Put in llms.txt (And What To Leave Out)

  • Include: canonical docs, policies, and “explainers” that you would cite yourself in an argument.

  • Leave out: duplicate pages, A/B test URLs, internal search results, and anything you blocked in robots.txt.

  • Update cadence: change it when you ship major docs restructures, rename key pages, or move URLs. Stale links make llms.txt worse than having none.

Advanced Setup: Markdown Alternates and Per-Section llms.txt Files

Stale links are the obvious failure mode, but advanced llms.txt setups often break for a quieter reason: the “best page” you list is a heavy HTML document that is hard for an agent to parse. The llms-txt v2 proposal suggests two upgrades for that: publish clean Markdown alternates for key pages, and scope llms.txt files by section so agents can choose the most relevant map.

Markdown Alternates (.md) for High-Signal Pages

The spec recommends offering a Markdown version of a page at a predictable URL. It gives two common patterns (llms-txt v2):

  • Append .md: /docs/page.html.md
  • Replace the extension with .md: /docs/page.md

For URLs that end in a slash (no filename), the spec says to append index.html.md or index.md. Example: /docs/ can map to /docs/index.md (llms-txt v2).

If you do this, keep the Markdown page plain: one H1, clear H2s, code blocks for commands, and minimal navigation boilerplate. Treat it like the page you would paste into a support ticket.

The spec also recommends advertising these alternates in HTML via link relations (llms-txt v2):

  • <link rel=“alternate” type=“text/markdown” href=“/docs/page.md”> points to the Markdown version.
  • <link rel=“describedby” href=“/llms.txt”> points to the llms.txt file that covers the page.

This does not guarantee any model will fetch the Markdown. It simply makes the relationship machine-readable.

Implementation note: in WordPress, you can add these tags with an SEO plugin that supports custom head markup or via wp_head. In frameworks like Next.js, you can emit them in <Head>. In Cloudflare Pages or Netlify, you can often add them in your layout template.

On large sites, use per-section files to reduce confusion. The spec allows /docs/llms.txt, which covers URLs under /docs/, and says that when multiple files apply, agents should use the most specific one (llms-txt v2). That “most specific wins” rule is a clean way to keep product docs, API docs, and help center content separated without turning one llms.txt into a dumping ground.

What Matters More Than llms.txt for AI Citations

A per-section llms.txt can keep your docs tidy, but it does not create citations by itself. If an AI system cannot crawl, index, or understand the pages you list, llms.txt becomes a directory that points to locked doors.

If you care about being quoted in AI answers, focus on the basics that control access and clarity. These basics also help Google, Bing, and any agent that behaves like a crawler.

  • Allow crawling in robots.txt. If you block /docs/ or your key product pages, an agent that respects robots.txt will not fetch them, even if you recommend them in llms.txt.
  • Make your “source of truth” pages indexable. Remove accidental noindex, avoid canonical tags that point elsewhere, and do not hide essential docs behind logins or aggressive bot checks.
  • Write pages that an LLM can quote cleanly. Use descriptive headings, short definitions near the top, and stable URLs for things like pricing, policies, and API auth.
  • Use internal links to show hierarchy. Link from “Getting Started” to deeper docs, and link back to the canonical explainer from related blog posts.

Make Your Pages Easy for Crawlers and Models to Parse

AI citations usually come from pages that look like good references. That means scannable structure and clear ownership of “the real answer.” Do this on the page before you worry about a separate llms.txt file.

  • Put the answer high on the page. For example, define a term in the first screenful, then expand with examples and edge cases.
  • Use consistent H1 and H2 headings. Agents often use headings as a table of contents when they chunk text.
  • Prefer HTML text over images and PDFs for anything you want quoted. If you must publish a PDF, also publish an HTML version with the same content.
  • Keep one canonical URL per concept. If you maintain multiple versions (docs, blog, help center), pick one as canonical and cross-link the others to it.

Then treat llms.txt as a pointer to those cleaned-up pages. It can help an agent choose the right URL faster, but it cannot compensate for blocked crawling, messy canonicals, or pages that bury the definition three scrolls down.

Check Whether AI Crawlers Can Reach Your Site (And Whether llms.txt Exists)

Screenshot of workspace Balzac

llms.txt only helps if an agent can actually fetch it and fetch the URLs you recommend. Before you spend time polishing an llms.txt example, confirm two basics: your site allows crawling where it matters, and /llms.txt (or your chosen subpath) returns a clean response.

Balzac’s free AI crawler checker is the fastest way to sanity-check both in one pass: https://hirebalzac.ai/free-seo-tools/ai-crawler-checker/. It reports whether /llms.txt exists and whether common AI crawlers can reach your pages.

  1. Enter your homepage URL (use the canonical version, including https).

  2. Run the check and review the fetch results for AI crawlers and for /llms.txt.

  3. Click through any flagged items and fix the specific cause (robots rules, auth walls, broken URLs, or server issues).

How to Interpret Common Results

If /llms.txt returns 200 OK, your server can serve the file. Next, open the file and spot-check a few listed URLs. Each one should also return 200 OK without requiring login, geo checks, or a cookie wall. If you publish per-section files like /docs/llms.txt, test those paths too.

If /llms.txt returns 404, you have not published it. That is not automatically “bad.” Google’s Lighthouse treats llms.txt as optional, so a 404 marks the audit as Not Applicable (N/A) (Chrome for Developers). If you do want the file, publish it at the root, for example https://example.com/llms.txt (same Lighthouse doc).

If you see a server error (5xx) when fetching /llms.txt, fix that before anything else. Lighthouse flags server errors when it tries to retrieve the file (Chrome for Developers). Common causes include misconfigured CDN rules (Cloudflare), origin timeouts, or an app route that intercepts /llms.txt and crashes.

If crawlers look blocked, inspect your robots.txt and any bot protection layers. Services like Cloudflare Bot Management can block or challenge automated fetches even when robots.txt allows them. Fix access first, then curate llms.txt. A reading list that points to blocked pages stays useless.

FAQ: llms.txt Example, Placement, and Common Pitfalls

Once your key pages are crawlable, llms.txt becomes a simple maintenance task: keep the file findable, keep the links accurate, and avoid recommending pages an agent cannot fetch.

FAQ

What is llms.txt?
An llms.txt file is a Markdown document, usually published at /llms.txt, that curates links to a site’s best “source of truth” pages for LLMs and AI agents. It follows the llms-txt v2 proposal, not an official web standard.

Where do I place llms.txt?
Put it at /llms.txt for site-wide coverage, or at a subpath like /docs/llms.txt to cover only that section. The proposal says an llms.txt file describes the pages under its path, so /docs/llms.txt covers /docs/, and when multiple files apply, the most specific one should win (llms-txt v2).

Can you show an llms.txt example?
Use the copy-paste llms.txt example earlier in this guide as a template. Keep it short: one H1, an optional blockquote summary, then a few H2 sections that group links with one-line notes.

What should I include in llms.txt?
Include pages you would cite yourself: getting started, core concepts, API reference, pricing, security, privacy policy, and a glossary. Favor stable URLs with clear headings and minimal UI clutter.

How often should I update llms.txt?
Update it whenever you move or rename a recommended URL, restructure docs navigation, change pricing or policy pages, or deprecate features. Treat it like a changelog for your “best sources,” not a set-and-forget file.

Will llms.txt improve rankings or guarantee AI citations?
No. No major AI search engine has promised to use llms.txt for rankings or citations. Ahrefs notes that no major LLM provider currently supports llms.txt, including OpenAI, Anthropic, and Google (Ahrefs).

What are the most common pitfalls?

  • Listing blocked URLs. If robots.txt disallows a path or Cloudflare Bot Management challenges requests, your llms.txt points to pages an agent cannot fetch.
  • Linking to thin pages. If the page does not answer a question clearly, the note in llms.txt will not save it.
  • Letting links rot. A stale llms.txt sends agents to 404s and outdated docs.
  • Dumping everything in. llms.txt works best as a curated shortlist, not a second sitemap.

How do I check whether /llms.txt exists and is reachable?
Fetch /llms.txt in a browser or with curl -I and confirm you get 200 OK. Fix access first, then curate the file.

Sources

← All articles

Free AI visibility check

Is ChatGPT citing your site?

See if ChatGPT search and Google AI Overviews cite you for your top searches, and who they cite instead. About a minute, no account.

Want content like this on autopilot?

Balzac researches, writes and publishes articles like this one to your site. Every week.

3 free articles · No card needed