TapCub is live — analytics, insights and live chat, free on one platform
AI crawlers

llms.txt: what to write, and how to check whether anyone reads it

llms.txt is an invitation, not a rule. A short guide to writing one that assistants actually use, and the ledger check that tells you whether they did.

TCTapCub Team Published Sep 17, 2026 7 min read AI crawlers

Try it on your own data

Every report in this article is in the free plan. One snippet, cookieless, up to 10 sites.

Start free

llms.txt is a plain-text file at the root of your site that tells AI assistants which pages you most want them to read and cite. It is young, informal and voluntary, which means two things: writing one is cheap, and nobody will tell you whether it worked. This guide covers what to put in the file, how it relates to robots.txt, and, the part most guides skip, how to verify from your own crawler ledger whether assistants fetched it and followed it.

What you will learn

  • What llms.txt is for and how it differs from robots.txt and a sitemap
  • A structure and template you can adapt in an hour
  • How to verify from the crawler ledger that the file and the pages it lists are actually fetched

What llms.txt is, and is not

Think of three files. A sitemap lists everything, for search engines, with no opinion. robots.txt sets permissions: which agents may fetch which paths. llms.txt is curation: the dozen or so pages that best explain who you are and what you offer, each with a one-line description, written for a reader that will summarise you to someone else. It does not grant or deny anything. A crawler you block in robots.txt will not read llms.txt, and a crawler you allow is free to ignore it.

Its value is in the moment an assistant is asked about your product or topic and has a few seconds to fetch something. A good llms.txt makes the right pages the obvious choice.

Because the format is informal, support varies by vendor and changes often. Treat the file as low-cost insurance: an hour to write, a few minutes a quarter to maintain, and a measurable effect you can read in your own logs rather than take on faith.

What to put in it

Keep it short and structured. A title line with your name, a one-paragraph description, then sections of links grouped by purpose: product pages, documentation, pricing, policies, and a few long-form articles that answer common questions well. Each link gets a short description that says what a reader will find, in plain words. Use absolute URLs. Avoid marketing adjectives; assistants quote descriptions, and “the leading platform” reads badly in someone else’s answer. Here is the one we publish, trimmed:

llms.txttrimmed example
# TapCub

> Cookieless web analytics, product insights and live chat on one snippet. Free plan, 7 SDKs, data stored by region.

## Product
- [Web analytics](https://tapcub.com/analytics): visitors, sources, pages, bots and performance for websites.
- [Insights](https://tapcub.com/insights): funnels, retention, paths, AARRR and attribution across 14 ad platforms.
- [Chat](https://tapcub.com/chat): live chat with proactive invites, AI replies and two-way translation.

## Pricing and policies
- [Pricing](https://tapcub.com/pricing): Free, Pro, VIP plans in USD.
- [Privacy](https://tapcub.com/privacy): what is collected and for how long.

## Guides
- [Why UV never matches](https://tapcub.com/blog/why-uv-never-matches): reconciling visitor counts between tools.
- [AI crawler ledger](https://tapcub.com/blog/ai-crawler-ledger): identifying GPTBot, ClaudeBot and others.

How it relates to robots.txt

The two files should agree. Every URL in llms.txt must be allowed for the agents you want to read it; a page listed in llms.txt and blocked in robots.txt is a contradiction the crawler resolves by not fetching. Conversely, the pages you block for training crawlers (user content, search results, logged-in areas) should not appear in llms.txt either. The table summarises what each file does, and the ledger article has a robots.txt that blocks training and allows answers.

FileAudienceSaysEnforced by
sitemap.xmlSearch enginesHere is everything, with change datesNothing; a hint for indexing
robots.txtAll crawlersThese agents may / may not fetch these pathsCrawler good behavior; edge rules for the rest
llms.txtAI assistantsThese are the pages worth reading and citingNothing; a curated invitation

Check whether anyone reads it

The honest answer to “does llms.txt work?” is “check your ledger”. TapCub’s crawler ledger records fetches of /llms.txt by agent, and for each AI agent shows what share of its fetched pages were listed in the file. A rising share after you publish is the signal. A flat share means the agent is not using it yet, which is useful to know before you spend more time on it.

The sample below is a month after publishing. PerplexityBot fetched the file weekly and 91% of its page fetches were listed pages; GPTBot fetched it once and stayed at 62%, roughly the share you would expect from a crawler that reads the sitemap instead.

app.tapcub.com/sites/demo/bots/llms-txt

/llms.txt fetched · 30 d

14

Agents that read it

4

Listed-page share · AI agents

74%

▲ 21 pts

/llms.txt · fetches and effect by agent

Published Aug 19
AgentFetched llms.txtLast fetchPage fetchesListed pagesTrend
PerplexityBot52 days ago24091%Rising
ChatGPT-User46 days ago18883%Rising
Claude-User39 days ago9679%Rising
GPTBot127 days ago1,20462%Flat
Googlebot112 days ago3,812—n/a
Sample data
Sample view one month after publishing. Answer agents read the file weekly; the training crawler read it once.

Keep it current

llms.txt rots like any index. Review it quarterly or whenever you launch a product page, retire a doc or change pricing. Keep the descriptions honest: if an assistant quotes a line about a feature you removed, that is on the file. Add new long-form articles sparingly; the file should stay readable in one screen. Put a dated comment at the top so you can see when it was last touched, and check the ledger a month after each change.

If your site has several languages, publish one file per language directory or one combined file with a section per language, and keep the descriptions in the language of the pages they point to. Assistants answer in the language of the question, and a description in the wrong language is the one most likely to be dropped.

Common mistakes

Four we see repeatedly. Listing every page (that is a sitemap; curate instead). Relative URLs or URLs that redirect (fetchers follow one hop at most). Descriptions written as slogans rather than summaries. And publishing the file without checking the ledger, then concluding it does not work. Write it in an hour, verify it in a month, and adjust. The AI crawler check tells you which agent a user agent string belongs to while you are reading the log.

TC

TapCub Team

The people who design, build and support TapCub. We write about what we measure on our own site and what customers ask us most.

Check whether assistants read your llms.txt

The bot view records every fetch of the file and shows, per agent, how much of what it reads is on your list.

See your data clearly. Find your growth.

Every click, backed by data. Install one line of code and see your first numbers in a minute.

No credit card · Free plans for analytics and chat · Cookieless analytics