Yoast SEO, Rank Math, AIOSEO, and SEOPress sold Schema.org as a ranking lever, but it’s only a rich-result format. They are making the same move with llms.txt. They’re selling you something that doesn’t work.
Over the past six months, we counted 4.4 million HTTPS requests on WordPress sites we operate. That includes this website. About 324,000 were AI hits — requests that name a known AI bot. Most of those hits are unattended crawls, not people asking ChatGPT about you. The full table is in the September 2026 report. We’ll remeasure if that pattern changes.
We found there is no cheat code for AI ranking. Ranking in AI is just SEO. You need to do a little more of the writing: complete answers on the pages they fetch. In this article, we break down what we found.
If a checklist tool highlights an AI feature, that check is bogus. PageSpeed Insights, Lighthouse, and SEO plugins will flag the absence of llms.txt, a GEO score, or WebMCP. But those checklists won’t rank your site.
Remember AMP? They also shoved that down our throats.
What are GEO and AEO?
GEO (Generative Engine Optimization) is a marketing name. So is AEO (Answer Engine Optimization). They are not standards.
Google Search says optimizing for its generative AI features is still traditional SEO. Those features sit on ordinary Search ranking and quality systems. AI Overviews and AI Mode are not a second ranking system.
ChatGPT Search and Claude Search use their own crawlers. A toggle labeled GEO does not make ChatGPT fetch your site.
AI Trainer, AI Search, and AI Agent
AI Trainer, AI Search, and AI Agents all work differently. Mixing the three is how “block ChatGPT” advice goes wrong.
An AI Trainer is a bot that crawls pages to train an LLM (large language model). It uses your page’s content. It does not store those URLs. It requests many URLs, often in the same second. That burst is a crawl, not a person reading. It also guesses extra paths and WordPress-ish slugs.
AI Search is the crawler that fills a searchable index for an assistant’s search product with URLs. That index is a database the product looks up when someone asks. It does not train the model. It requests many URLs, as a search engine does. OpenAI says its search crawler controls appearance in ChatGPT Search answers. Anthropic says its search crawler indexes content for Claude Search.
An AI Agent is a client that works on behalf of a user. It fetches a URL from a question, a paste, or a citation. If that page does not answer, it tries another URL. Most sessions still stop after one page.
| Provider | AI Trainer | AI Search | AI Agent |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Meta | Meta-ExternalAgent | Meta-WebIndexer | Meta-ExternalFetcher |
| SpaceXAI | — | grok-agent | |
| Microsoft | Bingbot | — | |
This table is not exhaustive. In our logs, Grok uses one name for search and fetch. Bingbot covers both training and search.
Do AI Agents honor robots.txt?
They do not treat the file the same way.
- OpenAI says robots.txt rules may not apply when a person asked ChatGPT to fetch.
- Perplexity says its user-invoked fetch generally ignores the file.
- Anthropic says Claude honors robots.txt when a person asked.
- A Chrome-like agent has no bot name, so robots.txt cannot name it.
To learn about blocking them, see our KB article on robots.txt blocks.
AI Trainers and AI Search hit ordinary WordPress sites whether anyone asked about them or not. AI Agents fetch URLs when people already want them. A GEO file does not attract AI Agents. An AI Trainer guessing a short slug is still an AI Trainer.
Do AI Agents start at your homepage?
No. An AI Agent starts on a specific page, not the homepage.
AI Trainers and AI Search can start at the homepage. They then request more pages. AI Trainers take the content. AI Search adds pages to its index.
An AI Agent looks further when the first page does not answer. In the traffic we measured, most sessions still stopped after one URL. The first page was enough.
We prompted various coding assistants the same questions about The SEO Framework. They fetched pricing, the refund policy, a knowledge base article, and our about us page. They never fetched the homepage. When a guess missed, they hit 404 and tried other slugs. Search didn’t send them to these pages first. They already knew the site from training data. For your site to be in that set of URLs, it still has to be found and indexed. That is SEO.
Landing pages are more important than ever. That includes the homepage, core product and pricing pages, knowledge base articles, and the about us page. People paste those URLs, ChatGPT Search and Claude Search cite them, and AI Agents guess them.
The llms.txt file
llms.txt is a proposed file at the root of your site. Plugins sell it as an AI ranking lever.
We first tested it in September 2025. We tested it again in September 2026.
Conclusion: llms.txt is not how ChatGPT, Claude, or Grok find or rank you.
Technology adoption scanners and SEO-tooling crawlers request the file. Those bots check who’s using it, or they collect site data for SEO reports. They do not rank you in search. In our logs, no AI Agent asked for it. Coding agents can still fetch the file when a person or tool already pointed them at a documentation site. That is not ChatGPT Search ranking your WordPress page. That is a specific user request.
Yoast SEO, Rank Math, AIOSEO, and SEOPress generate llms.txt. Their docs say it helps ChatGPT or “AI” understand the site and get you cited. AIOSEO turns it on by default—their plugin arrives on your site the same way: without consent.
Who requests the llms.txt file?
The public pitch for the file is llmstxt.org. That site is a proposal. Documentation platforms publish the file so tools already pointed at those docs can pull it. That’s not proof that ChatGPT Search fetches /llms.txt on your WordPress site.
Google Search ignores llms.txt. The file does not affect rankings, including AI Overviews and AI Mode.
On the TSF network, we saw about 270 GET /llms.txt requests. On ordinary WordPress sites on the same host, we saw about 300. Those hits were technology adoption scanners, not AI. Zero came from ChatGPT, Claude, or Grok.
This matches GitHub issue 732. After we placed a dummy file, the April 2026 report still showed zero AI-bot hits. Roots.io posted 57 days of logs: about 80,000 AI hits, 101 requests for the file, zero from AI bots.
Independent log studies land in the same place.
- Ahrefs measured about 38,000 valid files across 137,210 domains in May 2026. 97% of those files got zero requests. Of the existing traffic, AI retrieval bots accounted for about 1.1%.
- Limy counted 408
/llms.txtrequests in 515 million AI-bot events.
Should you ship the llms.txt file now?
There’s no early-mover advantage. The rush to ship llms.txt now is marketing, not strategy. We will implement it when an AI provider announces they consume it. A PageSpeed Insights note is not that announcement.
When that happens, they will fetch the file on your site that day. You do not need it sitting there beforehand.
An AI Trainer will not use llms.txt during training. It already fetches pages from links throughout your site. It trains on that content, not on a curated list of excerpts.
Why does PageSpeed Insights want the llms.txt file?
PageSpeed Insights only suggests. You do not need the file. PageSpeed Insights uses Lighthouse. A Lighthouse note is not Google Search.
Lighthouse now has an experimental Agentic Browsing category. One audit fetches /llms.txt. Chrome’s docs state that the file is optional. PageSpeed Insights shows that audit.
Google Search still ignores the file. AI Overviews and AI Mode do not use it.
You may get a red recommendation on your WordPress site anyway. That note is a lab check. It does not change rankings. It does not make an AI Agent fetch the file. The note may warn that models cannot train on your site. That is Lighthouse copy. It is not a Search ranking document.
Shipping a stub to silence the note can fail the audit. An empty file scores worse than no file.
We will not add a generator so a PageSpeed fraction turns green.
Use PageSpeed Insights to see how your site could be made faster. A slow site gets bounced more often. Some agents and embeds fail if your site responds too slowly. That is the job. The llms.txt audit is not. See bounce rate reduction.
Why does PageSpeed Insights annotate WebMCP?
WebMCP is a proposed Chrome API and PageSpeed is a Google product. It registers in-page tools so an agent in the browser can fill a form, book a table, or check out. It is not ranking.
A blog post, a knowledge base article, and most WordPress marketing pages have no tool to register. TSF will not add WebMCP attributes so a PageSpeed fraction turns green.
Should you serve Markdown to AI clients?
Some AI Agents send the HTTP header Accept: text/markdown. That header asks your site for a Markdown copy of the page instead of HTML. We keep a separate log of those requests on the TSF network (April through September 2026). This table lists only the requests that asked for Markdown.
| Who | Hits | What they’re doing |
|---|---|---|
| ShapBot (Parallel) | 3,813 | Crawler harvest, Markdown preferred. |
| Claude-User (Claude Code) | 1,833 | AI Agent, Markdown first. |
| grok-agent | 17 | AI Agent, Markdown first. |
| ChatGPT-User | 0 | Doesn’t negotiate Markdown in our logs. |
There’s a public pitch for Accept: text/markdown at acceptmarkdown.com. From our logs, coding agents send it. ChatGPT’s browse tool didn’t.
Google and Bing don’t send Accept: text/markdown. AIOSEO already serves a Markdown copy for “AI engines” at a .md URL. An agent that already fetched the HTML page still has to discover that extra endpoint. That is a second crawl. That spends more tokens, not fewer. SEOPress ships Markdown output too.
Serving Markdown to one client and HTML to the next requires Vary: Accept. That header poisons page caches. A poisoned cache serves Markdown to your visitors. Shortcodes, forms, and paywalls leak into that Markdown.
Agents already extract the page they fetch. They convert HTML on their side. They do not need a Markdown copy from your WordPress site. The converter depends on the request. The agent adjusts. TSF will not add this as a plugin feature.
Does semantic HTML make agents extract the article?
No. Google’s docs describe data-nosnippet for snippets and AI Overviews. Landmarks such as <main> and <article> help browsers and humans parse the page. They are not a documented Google extraction signal. The extractors we tested copied visible text. They ignored data-nosnippet and the landmarks. They also stripped scripts and styles, including JSON-LD that contains Schema.org.
We assumed the converters extract <main> and <article>. They do not. We pasted pages that already have both landmarks. Header, navigation, breadcrumbs, and a footer marked data-nosnippet still landed in the paste. All of that sits outside both landmarks.
The HTML is already correct. The built-in converters are the defect.
This is our prophecy: Agents ought, should, and will build better tooling around the semantic web. Put the complete answer in <article>. Well-written WordPress themes do this automatically. If the page has no landmark, converters have no article to take. That’s a document problem, not a missing llms.txt.
We found that self-hosted Elementor pages aren’t semantically valid HTML and may omit <main> and <article>. A post is one widget among other widgets. Our prophetic converter that wants “the article” will have no landmark to take.
Yet llms.txt keeps winning the debate: It’s easy to implement and forget about. The cost of the file is not that it might be useless. The cost is what it displaces. An hour on the file is an hour not spent improving your pages. Writing the page takes more time. It’s harder to show as progress. That’s why speculative tricks keep winning the meeting.
What should you do instead
- Ignore AI flags on checklist tools. PageSpeed Insights, Lighthouse, and SEO plugins will highlight
llms.txtand GEO scores. Those checks are bogus. See why PageSpeed Insights wants the llms.txt file. - Write the landing pages. Homepage, core product pages, knowledge base articles, About Us, and pricing get the fetches. Put the complete answer on each. An AI Agent arrives with a URL from search, a paste, a citation, or a guessed WordPress-ish slug. The page it fetches is the document. See practical SEO tips.
- One canonical URL. Search follows the canonical. When an AI Agent arrives from search, that is the URL it fetches. Don’t invent a second URL for “AI.” We have honed canonical URLs for over a decade. See advanced query protection.
- Landmarks. Put the article in
<main>/<article>. Mark header, nav, and footerdata-nosnippetif they should stay out of snippets. See whether semantic HTML makes agents extract the article. - Structured data. Keep Schema.org for Google Search. The extractors we tested stripped scripts and styles, including JSON-LD. That is not a GEO input. See structured data supported by The SEO Framework.
- robots.txt. The AI crawler block in SEO Settings → Robots Settings → Robots.txt is optional. Use it if you want AI Trainers off the site. That block does not stop ChatGPT Search or ChatGPT’s fetcher. See robots.txt blocks.
Will The SEO Framework add AI SEO features?
No.
| Feature | Why not |
|---|---|
| GEO score or AEO checklist | Those labels are marketing names. On Google Search, they still mean traditional SEO. |
| llms.txt | AI Agents don’t request it. Google Search ignores it. We’ll add it when a provider announces they consume it. See whether llms.txt gets you cited. |
| Markdown for AI clients | From our logs, ChatGPT’s browse tool didn’t ask. Serving both formats poisons page caches. See whether you should serve Markdown. |
| WebMCP or a PageSpeed Agentic score | A lab fraction is not a ranking factor. See why PageSpeed Insights wants llms.txt. |
An “SEO score” teaches you how to chase the score, not how to write the page. We skipped readability scores for the same reason.
TSF still ships robots.txt, canonicals, redirects, structured data, and clean HTML. Those are the bits that provide a clean public URL to whoever fetches it.
What might change later?
This section is speculation. It may be wrong next week.
An AI provider may one day announce they consume llms.txt. That announcement is the signal. A plugin generator, a Google AI-readiness page, and a Lighthouse note are not.
Until then, write the complete answer on the public WordPress page. An AI Agent fetches that URL. If the page answers, the agent will use that.