llms.txt Validator and Generator: Free Spec Checker
Enter a domain or paste an llms.txt file. The tool checks the format line by line against the llmstxt.org spec, points to the line with each problem and shows how to fix it. Switch to Generate to build a new llms.txt from a sitemap or a page list. Pasted text is processed in your browser and never uploaded.
Key takeaways
- The tool checks llms.txt line by line against the llmstxt.org spec and points to the line with each problem; the Generate tab builds a new llms.txt from a sitemap or a page list.
- Google Search does not use llms.txt, and there is no evidence yet that it raises citations in ChatGPT, Perplexity or similar products.
- Developer docs, APIs and SaaS documentation are worth an llms.txt; for export stores, e-commerce and content sites it is a low priority.
- Pasted text is processed in your browser and never uploaded.
What is llms.txt?
llms.txt is a Markdown file at the root of a website (https://yourdomain/llms.txt). Jeremy Howard of Answer.AI proposed it in September 2024 and published v2 of the proposal in August 2026. It gives AI agents a concise index of the site: what the site is, which pages are worth reading, and what each page covers. An agent helping a user with research or code can read this index first and then open the links it needs, instead of stripping navigation and ads out of HTML pages.
It does a different job from robots.txt and sitemap.xml:
- robots.txt says which crawlers may fetch which paths. It is access control. To opt out of AI training you write rules in robots.txt; llms.txt cannot do that.
- sitemap.xml lists every URL you want indexed, for search engines.
- llms.txt is a curated index of the important pages, each with a one-line description, read by AI agents on demand while answering a question.
The spec defines the file in this order: an optional BOM; an H1 with the site or project name (the only required part); a blockquote summary; any number of paragraphs or lists with details (no headings); and any number of sections delimited by H2 headings, each a list of links in the form - [name](url): notes. A section named Optional holds secondary links that can be skipped.
The evidence: does llms.txt do anything today?
Before making one you should know what it does in practice. The table uses the evidence levels used across this site (A Confirmed, B Documented, C Patented, D Speculative or third-party data).
| Claim | Level | Source |
|---|---|---|
| Google Search does not use llms.txt; having one neither helps nor hurts | A | Google Search Central: optimizing for generative AI features (updated 2026-07-10) |
| John Mueller compared llms.txt to the keywords meta tag and said server logs show AI services do not even request it | A | Search Engine Journal (April 2025, originally on Reddit) |
| Gary Illyes said at Search Central Live that Google does not support llms.txt and has no plans to | A | Search Engine Land (July 2025) |
| Chrome Lighthouse 13 added an Agentic Browsing category with an audit for a valid llms.txt | A | Lighthouse source |
| Ahrefs looked at 137,210 domains: 28% had llms.txt, and 97% of those files got zero requests in May 2026; only 19.5% of the requests that did arrive came from AI tools | D | Ahrefs study (June 2026) |
| SE Ranking analysed about 300,000 domains and found no link between having llms.txt and how often a domain is cited by AI | D | Search Engine Journal |
In short, llms.txt does not affect Google rankings or AI Overviews, and there is no evidence yet that it raises citations in ChatGPT, Perplexity or similar products. Google is not of one mind: the Search team says it is unnecessary, while the Chrome team includes it in Lighthouse’s agent-readiness checks. The spec’s author notes in v2 that OpenAI, Anthropic and Google Gemini publish llms.txt files for their own developer docs, and that coding agents read such files.
For experts: details of the Ahrefs data
- The data comes from server logs and live traffic of Ahrefs Web Analytics customers, screened for soft 404s and “phantom” files that actually return the home page.
- Of the 3% of files that got any requests, 96% of requests came from bots and 4% from people. The largest bot group was SEO audit tools (21.7%), followed by unidentified bots, general crawlers and technology profilers.
- AI-related requests totalled 19.5%: agents 10.5% (Claude-Code first), training crawlers 5.3% (GPTBot first), AI assistants 2.5% and AI retrieval bots only 1.1%.
- Sites without llms.txt received no AI bot requests for it, so “AI goes looking for the file” cannot be used as evidence that it helps.
Should you make one?
It depends on the kind of site:
- Developer docs, APIs, SaaS documentation: worth doing. Coding agents read these files, and documentation platforms such as Mintlify can generate them. Also serve
.mdversions of important pages. - E-commerce, content and lead-gen sites: low priority. It is quick to make and does no harm, but do not expect rankings or AI citations from it. Spend the time first on what Google says matters: crawlable, indexable pages with content that adds something new.
- Sites that already have one: no need to remove it. Use this tool to make sure the format is right, then use the AI Crawler Log Analyzer to see whether any AI crawler has requested
/llms.txtin your server logs, and decide from your own data whether to keep maintaining it.
How to use the tool
- Under Validate, enter a domain and this site’s server fetches
/llms.txtand/llms-full.txt, or paste the file. With no file at hand, click “Example (with common mistakes)” to see what the tool does, or “This site’s llms.txt” for a real file. - Start with the hard requirements: the H1 the spec requires and the checks of Chrome Lighthouse. Each one passes or fails.
- Go through errors, warnings and tips under Issues. Clicking a line number jumps to that line in the editor; “Why · How to fix” explains the problem and gives a fix you can copy. Bullets, full-width colons, relative URLs and missing spaces after
#can be fixed in one click with Auto-fix, and undone. - Structure shows what a program extracts from the file; Links shows duplicates and external links; File info shows status code, Content-Type, encoding and llms-full.txt.
- If you have no llms.txt, switch to Generate, paste a sitemap.xml or a page list, arrange the groups, add notes, then copy or download the result.
What the tool checks, and why
Checks come from three kinds of source, and severity and evidence level are shown separately:
- The spec text (A): requirements the llmstxt.org proposal states, such as the required H1, no headings between the H1 and the first H2, and a
[name](url)link in every list item of an H2 section. Breaking these is an error or a warning. - Reference implementation behaviour (B): the spec’s author publishes a parsing function in the
llms_txtPython package, and the tool reproduces how it reads a file. It only accepts link lines starting with-and raises an error on*bullets, numbered lists and plain text inside sections; it keeps only the last of two sections with the same name; and it treats###as a new section. These are not necessarily spec violations, but they show that some programs will misread the file. - Our advice (D): suggestions from experience with no official basis, such as writing notes for links, keeping the file small, not listing hundreds of links, and using full URLs. These are tips, never errors.
The Chrome Lighthouse llms.txt audit (A) looks at four things only: a 5xx response for /llms.txt fails, a 4xx is not applicable; the file has an H1 written as “#, space, text”; it has at least one [text](url) link; and it is at least 50 characters long. The hard requirements panel shows each of them.
The tool gives no overall score. A single number would mix “missing H1” with “a few links without notes” and hide what to fix first.
Common mistakes and fixes
Text that is not a link inside an H2 section. For example a sentence of explanation, or key-value items like - Name: Jane. Sections are link lists in the spec, and the reference parser fails on such lines. Put prose between the H1 and the first H2, or after a link as its notes:
1 | ## Policies |
A dash instead of a colon before the notes. In - [Tents](https://...) - Family tents the dash is not a separator and the notes are lost. Use a colon and a space, : . On Chinese sites the same happens with the full-width colon “:”.
Relative URLs. - [Lanterns](/collections/lanterns) has no host when the file is read on its own. Write https://www.example.com/collections/lanterns.
An optional section not named Optional. “Optional links” or “Secondary” is not recognised by tools that follow the convention. Since v2 Optional is only a convention with no mechanical meaning, but naming it exactly ## Optional costs nothing.
/llms.txt returns a web page. Single-page apps often route every path to the home page, and some 404 pages return 200. Open /llms.txt in a browser; if you see a page of your site, put a real file in the root and make sure the server returns it directly.
Garbled characters. The file is not UTF-8, or the Content-Type lacks charset=utf-8. Re-save the file as UTF-8 and serve it with Content-Type: text/plain; charset=utf-8. “Why · How to fix” includes Nginx, Apache and Vercel examples.
For experts: edge cases of the reference parser
parse_llms_filesplits sections with^##\s*(.*?)$, so a###line also becomes a section whose name starts with#.- Every non-empty line in a section must match
-\s*\[([^\]]+)\]\(([^)]+)\)(?::\s*(.*))?; a single non-matching line raises an exception and the whole file fails. Escaped brackets\]in a link name also fail to match, so the generator turns brackets into parentheses. - The head is matched with
^#\s*(.+?)$\n+(?:^>\s*(.+?)$)?\n+(.*). When there is nothing but the H1 and the summary before the first H2, this either fails (an exception) or reads the summary as body text. A paragraph after the summary avoids both. - At the bottom of Structure you can expand the JSON the reference parser produces and compare it with the structure the spec describes.
Using the generator
The generator accepts four kinds of input: a full sitemap.xml; one URL per line; lines of “title | url | notes”; and a Markdown link list. Imported URLs are grouped by their first folder, so /collections/ and /blogs/ become two groups. Link titles are derived from the last part of the URL, with non-ASCII URLs decoded.
You can then rename groups, drag links to reorder them or move them to another group (the menu and arrow buttons beside each link do the same with a keyboard or on a phone), and put secondary pages in Optional. The output on the right updates as you go; copy it, download it, or send it to Validate for another check.
The generator does not use AI to write notes. Write one sentence per link on what the page holds. llms.txt is a curated index, not a copy of the sitemap, so keep a few dozen of the most important pages. If you paste a sitemap index (sitemapindex), the tool lists the child sitemaps; open one and paste it instead.
FAQ
Does llms.txt affect Google rankings?
No. Google’s documentation states that Google Search, including AI Overviews and AI Mode, does not use llms.txt, and that having one neither helps nor hurts.
Can llms.txt stop AI from training on my content?
No. llms.txt is an index for agents, not access control. To block training crawlers such as GPTBot and ClaudeBot, write rules in robots.txt and check them with the AI Crawler robots.txt Checker.
Does llms.txt have to be at the root?
v2 of the proposal allows any path: a file covers the pages under its path, and the most specific file applies. For example /docs/llms.txt describes only pages under /docs/. A multilingual site can put one at the root and one under /en/, as this site does. When you enter a domain the tool fetches only the root /llms.txt; paste files from other paths.
What is llms-full.txt, and do I need one?
llms-full.txt is not part of the llmstxt.org proposal. It is a convention from some documentation platforms that puts the full text of the site in one file. v2 of the proposal recommends .md versions of important pages instead. When you enter a domain the tool also fetches /llms-full.txt and shows whether it exists and its size; a button loads it into the editor for checking.
Does the tool upload my content or crawl my site?
Pasted text stays in your browser. When you enter a domain, this site’s server requests only two fixed addresses, /llms.txt and /llms-full.txt, and returns only the status code, Content-Type, size and the first 500 KB of text; if the answer is an HTML page, no page content is returned. Fetches are rate limited (30 per IP per hour, and one check uses two), and the same URL is served from a 10-minute cache. The tool does not visit the links in the file; the Links tab has Open buttons so you can spot-check them yourself.
Why does this site’s own llms.txt show warnings?
The Author section of this site’s llms.txt uses key-value items such as “- Name: Linus Li”, which do not fit what the spec expects in an H2 section, and the reference parser cannot read it. Click “This site’s llms.txt” to see the issues.
Why is there no overall score?
A score mixes problems of very different weight and suggests that a high number means AI will cite you. The tool shows the spec requirement and the Lighthouse checks as pass or fail, grades everything else as errors, warnings and tips, and gives the source of every rule.
Further reading
- AI Crawler Log Analyzer: check in your server logs whether AI crawlers read your llms.txt
- AI Crawler robots.txt Checker: see whether GPTBot, ClaudeBot and others can crawl your pages
- Schema Visualizer: check your structured data
- The llmstxt.org proposal
- More free SEO and GEO tools