Tabkeel
September 28, 2026·Francisco Ferreira·13 min read

How to Create an llms.txt File for an AI-Built Site

Quick answer

To create an llms.txt file, write a plain Markdown file named exactly llms.txt, put your product name as an H1, add a one-line summary in a blockquote, then list your most important pages as Markdown links under H2 headings. Save it to your site root so it loads at yourdomain.com/llms.txt. On a Next.js or Vite build, that means dropping it in the public/ folder. The whole thing takes about ten minutes, and the catch is below.

Here is the part every other guide skips. An llms.txt file points AI models at your key pages, but if those pages render only in the browser, the model follows the link and reads an empty shell. So the file works exactly as well as the pages it lists, which on an AI-built site is often not at all. Learn how to create an llms.txt file the right way and you fix the file and the pages together. Do it the WordPress-plugin way most tutorials teach and you ship a tidy signpost pointing at rooms with no furniture in them.

16,670domains publishing an llms.txt by September 2026, up from 951 in July 2025
~17xgrowth in adoption in 14 months, faster than any confirmed proof it works
0major AI platforms that have confirmed they read the file

That tension, real momentum against zero confirmation, is why this guide treats llms.txt as a cheap bet worth placing, not a growth lever worth believing in. You will get the exact steps, a copy-paste example, where the file goes on a founder's stack, and an honest read on whether it does anything yet.

What is llms.txt, in one sentence

An llms.txt file is a Markdown file at your domain root that gives large language models a curated map of your most important pages, so a model with a small context window can find your real content instead of parsing a cluttered site. Jeremy Howard, co-founder of Answer.AI, proposed the format in September 2024. The idea borrows from robots.txt but inverts the intent: robots.txt tells crawlers what to skip, llms.txt tells them what matters.

It is one file, not a system. No plugin required, no build step, no API. You write it, you host it, you keep it current. And it lives next to two relatives worth naming now so you do not confuse them later.

FileWhat it doesWho it talks to
llms.txtLists your key pages as a curated indexLLMs looking for your best content
llms-full.txtBundles the full text of those pages into one fileLLMs that want the content inline, no extra fetch
robots.txtAllows or blocks crawler access, by path and user agentSearch and AI crawlers deciding what to fetch

Most founders need llms.txt and can ignore llms-full.txt until they have real documentation to bundle. The one that actually gates whether AI reads you at all is robots.txt, covered in the pillar on making your website readable to AI, because a blocked GPTBot never reaches either of the other two files.

The trap: llms.txt cannot fix a site the AI cannot read

Read this section before you write a single line of the file, because it decides whether the file does anything. If you built your marketing site with Lovable, v0, Cursor or a plain Vite setup, your pages probably render client-side. The server sends a near-empty body, the browser fills it in with JavaScript, and a human sees a finished page. Most AI crawlers do not run JavaScript, so they see the empty version.

What your llms.txt promises

Here is my pricing page, my product overview, my docs. Go read them, they explain everything.

What GPTBot receives when it follows the link

An HTML body containing one empty div and a script tag. No prices, no product copy, nothing to quote.

An llms.txt file is a table of contents. If the chapters are blank pages, a better table of contents changes nothing. This is the single most common reason an AI-built site stays invisible to models, and it sits upstream of every GEO tactic you will read about this year. So the honest order of operations is: confirm your key pages serve real HTML first, then write the llms.txt that points at them.

Checking takes thirty seconds. Open your pricing page, then use View Source, not Inspect. Inspect shows the rendered tree after JavaScript runs; View Source shows the raw HTML a crawler actually gets. If your prices and product description are missing from the source, that is your real project, and the llms.txt can wait an afternoon. The AI readability check runs that same test across your site and names the pages that come back empty.

How to create an llms.txt file, step by step

Once your pages serve real HTML, the file itself is quick. Five steps, worst mistakes flagged at each one.

  1. Pick the 10 to 20 pages that define your product. Homepage, pricing, product or features, docs, key blog posts, an about page. Not every URL you own. The whole value of the file is that it is curated, so a link to a thin or duplicate page dilutes it.
  2. Write a one-line summary of what you are. This goes in a blockquote right under the title and it is the line a model is most likely to lift verbatim. Make it a plain sentence a stranger would understand, not a tagline. "Tabkeel is an SEO and AI-readability checker for founders who built their SaaS with AI tools" beats "Ship with confidence."
  3. Structure it in Markdown. H1 with your product name, the blockquote summary, then H2 sections grouping your links. Each link is a Markdown link with a short description after it, so the model knows what each page is before fetching it.
  4. Save it as llms.txt, exactly. Lowercase, no capital L, no .md extension. A file named LLMs.txt or llms.md will not be found at the path models look for.
  5. Host it at your root. It must load at yourdomain.com/llms.txt, not in a subfolder. The next section covers exactly where that file goes on your stack.

Here is a real-shaped example for a small SaaS. Copy the structure, swap the content.

# Tabkeel

> Tabkeel is an SEO and AI-readability checker for founders who built their SaaS with AI tools. Paste a URL, get findings with evidence and a fix written as a prompt.

## Core pages
- [What Tabkeel checks](https://tabkeel.com/methodology): the seven fronts, from AI readability to billing integrity.
- [Pricing](https://tabkeel.com/pricing): Free, Founder, Studio and Agency plans with page and site limits.
- [Run the exam](https://tabkeel.com/check): paste a URL, no signup, get every finding.

## Guides
- [Make your website readable to AI](https://tabkeel.com/blog/make-your-website-readable-to-ai): why crawlers see an empty page and how to fix it.
- [Get cited by ChatGPT](https://tabkeel.com/blog/get-cited-by-chatgpt): what answer engines quote and how to earn it.

## About
- [About Tabkeel](https://tabkeel.com/about): who built it and why.

Prefer not to hand-write it? An llms.txt generator like Wordlift or Firecrawl will crawl your site and draft the file for you. Treat the output as a first draft, not a final answer, because a generator will happily list thin pages and miss your best blog post. The curation is the part only you can do well.

Where the file goes on your stack

This is the step that trips up founders on a modern build, because "put it at the root" means something different depending on how your site is served. The path in your repo is not the URL. Here is the mapping for the stacks AI-built SaaS actually ships on.

StackWhere the file goesResult URL
Next.js (App or Pages Router)Drop llms.txt in the public/ folderyourdomain.com/llms.txt
Vite / React (Lovable, v0 default)Drop it in public/, it is copied to the build rootyourdomain.com/llms.txt
Vercel static / plain HTMLPut it in the deployed root directory next to index.htmlyourdomain.com/llms.txt
WordPressUpload to the root via SFTP, or use a plugin toggleyourdomain.com/llms.txt

For most founders reading this, the answer is the first two rows: the public/ folder. Anything in public/ is served as-is from the root, so a file at public/llms.txt becomes yourdomain.com/llms.txt with no routing code. If you want the file generated from your content instead of maintained by hand, a Next.js route handler at app/llms.txt/route.ts can build it on the fly, but that is an optimization, not a requirement. Ship the static file first.

If your llms.txt still 404s after you deploy, the file is almost always in src/ or a component folder instead of public/. Assets that need to be served at a literal URL go in public/; anything the bundler processes does not. That one distinction fixes most missing-file cases on a Vite or Next stack.

Does anyone actually read it? An honest answer

No major AI platform has confirmed it uses llms.txt. Not OpenAI, not Google, not Anthropic. In 2025, Google's John Mueller compared it to the keywords meta tag, the SEO relic nobody trusts anymore, and Google has since said its systems do not use llms.txt for AI Overviews. So the file that 16,670 domains rushed to publish has no confirmed reader among the models it was built for.

So why publish it at all?

Because the cost is ten minutes and the downside is zero, while the upside is real if adoption catches up. Some AI-native tools and smaller crawlers already look for the file, and having a clean, current index of your best pages helps you regardless of who reads it. Publish it as a cheap hedge. Do not report it to your cofounder as a growth channel, and do not let it distract from the readability work that models demonstrably do act on.

Put plainly: llms.txt is a low-severity, low-cost signal. In Tabkeel's exam it sits near the bottom of the AI-readability front for exactly this reason, well below server-rendered HTML and structured data, which change whether a model can read and quote you at all. If you only have an hour, spend fifty-five minutes on readability and five on this file.

How to check your llms.txt is live

Two checks confirm the file is doing the one job you can verify. First, that it is served. Second, that it is served as plain text, not swallowed by your app's catch-all route.

Load yourdomain.com/llms.txt in a browser. You should see your raw Markdown, not your site's 404 page and not your homepage. If a single-page app returns its shell for the path, you have the soft-404 problem that also breaks real error pages, which is its own finding in the pre-launch checklist for AI-built sites. From the command line, confirm the status and type in one shot:

curl -sI https://yourdomain.com/llms.txt

You want a 200 status and a content-type of text/plain or text/markdown. An HTML content-type means your framework is rendering the path through your app instead of serving the file, which is the same routing mistake as the missing-file case above. Here is the honest limit of this check: it proves the file exists and is readable, and nothing more. No tool can currently confirm that ChatGPT or Claude ingested it, because none of them report that they did.

The mistakes that make the file pointless

Almost every wasted llms.txt fails in one of these ways. Read the list against your own file before you call it done.

  • Listing pages that render blank to crawlers. The trap from earlier, and the most damaging. A curated index of unreadable pages is still unreadable. Fix the pages, then list them.
  • Dumping every URL you own. A sitemap already does that, and it is the opposite of what this file is for. Twenty great pages beat two hundred mediocre ones.
  • Writing a tagline as the summary. The blockquote is the line a model quotes. "The future of shipping" tells it nothing. State what you are and who it is for.
  • Letting it rot. You rename a plan, kill a feature, move a page, and the llms.txt still points at the old world. A stale index is worse than none, because it feeds a model wrong facts about you.
  • Putting it in the wrong folder. A file in src/ instead of public/ never reaches the root URL. Deploy, then load the URL to confirm.

Reconcile the file with your live pages the same way you would reconcile your pricing page with your terms: read them side by side and make them agree. When plans or pages change, the index has to change with them, or an answer engine will confidently repeat a price you retired. That drift, between what your site says and what a model believes, is the exact thing Tabkeel's exam watches for across every front.

Ship the file, then check the pages it points at

Writing the llms.txt is the easy ten minutes. The work that moves the needle is making sure the pages it links actually reach a crawler, because a model can only quote what it can read. So publish the file today as a cheap hedge, then spend the real effort confirming your key pages serve their content in the HTML. Want to see which of your pages come back empty to an AI crawler before you bother indexing them? Point the Tabkeel exam at your URL and it returns the readability findings with the evidence attached and each fix written as a prompt you paste into your own builder, the most severe one shown in full at no cost. The deeper method behind why an empty shell voids every other tactic is in the guide on making your site readable to AI, and if the goal is getting quoted rather than just crawled, that is a separate measurement covered in getting cited by ChatGPT.

Frequently asked questions

How do I create an llms.txt file?

Write a plain Markdown file named exactly llms.txt: an H1 with your product name, a one-sentence summary in a blockquote, then H2 sections listing your 10 to 20 most important pages as Markdown links with short descriptions. Save it to your site root so it loads at yourdomain.com/llms.txt. On a Next.js or Vite build, that means the public/ folder. Before you list a page, confirm it serves real HTML to a crawler, or the file points at content the model cannot read.

Where do I put the llms.txt file on Next.js or Vite?

In the public/ folder. Anything in public/ is served from the root unchanged, so public/llms.txt becomes yourdomain.com/llms.txt with no routing code. If the file 404s after deploy, it is almost always sitting in src/ or a component folder instead. On Next.js you can also generate it from a route handler at app/llms.txt/route.ts, but the static file is the simpler first move.

Do ChatGPT and Google actually use llms.txt?

No platform has confirmed it. OpenAI, Google and Anthropic have not said they read llms.txt, and Google has stated its systems do not use it for AI Overviews. Adoption grew from 951 domains in July 2025 to 16,670 by September 2026, but that is publishers betting on it, not proof it works. Treat it as a cheap hedge, and put your real effort into server-rendered HTML and structured data, which models demonstrably act on.

What is the difference between llms.txt and robots.txt?

robots.txt tells crawlers what not to access; llms.txt tells LLMs which pages matter most. They solve opposite problems and do not replace each other. robots.txt is the gate that decides whether an AI crawler can reach your site at all, so a blocked crawler never sees your llms.txt. Get robots.txt right first, then use llms.txt to curate what a model finds once it is allowed in.

Can I use an llms.txt generator instead of writing it by hand?

Yes, tools like Wordlift and Firecrawl crawl your site and draft the file. Treat the output as a starting point, because a generator lists pages by what it can crawl, not by what represents you best, so it tends to include thin pages and miss your strongest content. The curation and the one-line summary are the parts worth doing yourself.

FF
Francisco Ferreira
Builds Tabkeel and runs the exam on AI-built sites every day. About

See what AI and Google read on your site

The exam crawls your public site and returns every finding with the evidence and a paste-ready fix. Free, no signup.

Run the free exam

More articles

← All articles