We Built an llms.txt Generator. Here's What It Won't Tell You
A free tool that generates a real llms.txt from your sitemap - plus an honest look at what the file is for, what it isn't, and which AI crawlers actually read it today.
Of the three free tools we shipped this week, this is the one most likely to get overclaimed by someone else's marketing page. So before describing what our llms.txt Generator does, it's worth being precise about what the file itself is and isn't.
What llms.txt actually is
It's a proposed convention - not a standard anything has to implement - for giving an AI system a curated, markdown-formatted map of a site's important pages at /llms.txt. Think of it as a sitemap written for a language model to read: an H1 with the site name, a one-line summary, then sections of markdown links grouped by area, each with a short description.
That's it. It's a plain text file. Nothing about it is enforced, verified, or required. A generated file looks roughly like this:
# Acme Widgets
> Acme sells custom widgets for makers and hobbyists.
## Docs
- [Getting Started](https://acme.example/docs/getting-started): Install the SDK and run your first widget.
- [API Reference](https://acme.example/docs/api)
## Blog
- [Why We Rebuilt Our Checkout](https://acme.example/blog/checkout-rebuild): What broke, what we learned.
Notice the API Reference entry has no description after it - that page didn't have a meta description, so we left it out rather than write one. More on why below.
What it is not
It is not robots.txt. robots.txt controls crawl access - what a bot is and isn't allowed to fetch. llms.txt controls nothing; a crawler that ignores it entirely loses no functionality, because it was never a gate on anything. Our Crawler Access Checker tests real access control. This tool does not, and conflating the two is exactly the kind of mistake that makes a technical reader stop trusting the rest of the post.
It is also not a ranking or citation signal. We have no evidence - and haven't seen credible evidence from anyone else - that publishing an llms.txt changes whether a page gets cited in an AI answer. [TK: stat - if/when we can measure citation rate correlation with llms.txt presence across our own audited sites]
Which crawlers actually read it
Honestly: adoption is uneven, and it's changing month to month. Some AI products reference llms.txt-style files when present; most major crawlers don't fetch it as a routine part of indexing today. We're not going to pretend otherwise to make this tool sound more useful than it is. Publishing one costs nothing, so there's little downside to having a correct one ready - but "little downside" is a different claim than "you should expect a return."
What the generator actually does
Given a domain, it:
- Discovers your sitemap - checking robots.txt's
Sitemap:directive first, then common paths, handling sitemap-index files and gzip. - Caps at the 200 shallowest URLs on larger sites, and says so in the output if it had to truncate.
- Fetches each page's real
<title>and meta description, concurrently, inside a 60-second budget - partial results are labelled as partial, never silently presented as complete. - Groups pages into sections by path (
/docs/,/blog/,/features/) and renders straight into valid llms.txt markdown.
Nowhere in that pipeline is there a model call. Where a page has no meta description, the generated entry just doesn't have one - we do not write filler copy to make the file look more complete than the site actually is. If you already have an llms.txt, we fetch and diff against it so you can see exactly what changed.
Sitemap discovery turned out to have more edge cases than the one-line description suggests. Some sites don't have a plain sitemap.xml at all - they have a sitemap_index.xml that points at a dozen smaller sitemaps, one per content type, and we have to fetch and merge across all of them. Some serve those files gzipped without necessarily saying so cleanly in the Content-Type header, so we sniff the actual gzip magic bytes rather than trust the header alone. None of this is exotic, but it's the kind of thing that makes "just parse the sitemap" a bigger job than it sounds.
Long-time readers of the free audit post will recognize one of the five structural checks it runs: whether an llms.txt, llm.txt, or ai.txt file exists at all. That check only ever told you yes or no. This tool is what you run once the answer is no and you want an actual file, not just a missing-file flag.
What you should actually do with it
What you should actually do with it
Treat the output as a strong first draft, not a finished file. Read it, edit the summary line, fix any section groupings that don't match how you'd actually organize your site, and then decide whether publishing it is worth the (small, but real) maintenance cost of keeping a second sitemap-like file in sync as your site changes.
Try it
getbotrank.com/tools/llms-txt - enter a domain and an email, get a real draft in under a minute.