The SEO Framework · KB

★︎ Start with TSF
  • Extensions
  • Documentation
  • Pricing
  1. Home
  2. Knowledge Base
  3. AI SEO Myths and Facts

AI SEO Myths and Facts — Contents

  • What are GEO and AEO?
  • AI Trainer, Search, and Agent
    • Do AI Agents honor robots.txt?
  • Do AI Agents start at your homepage?
  • The llms.txt file
    • Who requests llms.txt?
    • Should you ship llms.txt now?
    • Would llms.txt help if providers consumed it?
    • Why does PageSpeed want llms.txt?
      • Why does PageSpeed annotate WebMCP?
  • Should you serve Markdown?
  • Do agents extract semantic HTML?
  • What should you do instead?
  • Will TSF add AI SEO features?
  • What might change later?

AI SEO Myths and Facts

Published on September 15, 2026
Revised on September 23, 2026

Yoast SEO, Rank Math, AIOSEO, and SEOPress sold Schema.org as a ranking lever. Yet Schema.org is only a rich-result format. They are making the same move with llms.txt. They’re selling you something that does not work.

We ran two six-month log trials over twelve months on WordPress sites we operate. That includes this website. The latest trial counted 4.4 million HTTPS requests. About 324,000 were AI hits — requests that name a known AI bot. Most of those hits come from unattended crawls, not people asking ChatGPT about your site. The September 2026 report lists the full table. We’ll remeasure if that pattern changes.

There is no cheat code for AI ranking. Ranking in AI is just SEO. You still have to do a little more of the writing: serve complete answers on the pages AI Agents fetch.

If a checklist tool highlights an AI feature, that check is bogus. PageSpeed Insights, Lighthouse, and SEO plugins will flag the absence of llms.txt, a GEO score, or WebMCP. Those checklists will not rank your site.

Remember AMP? Google also shoved that down our throats.

What are GEO and AEO?

GEO (Generative Engine Optimization) is a marketing name. So is AEO (Answer Engine Optimization). They are not standards.

Google Search says optimizing for its generative AI features is still traditional SEO. Those features sit on ordinary Search ranking and quality systems. AI Overviews and AI Mode are not a second ranking system.

ChatGPT Search and Claude Search use their own crawlers. A toggle labeled GEO does not make ChatGPT fetch your site.

AI Trainer, AI Search, and AI Agent

AI Trainer, AI Search, and AI Agent all work differently. Mixing the three is how “block ChatGPT” advice goes wrong.

An AI Trainer is a bot that crawls pages to train an LLM. It uses your page’s content. It does not store those URLs. It requests many URLs, often in the same second. That burst is a crawl, not a person reading. It also guesses extra paths and WordPress-ish slugs.

AI Search is the crawler that fills an assistant’s searchable index with URLs. The assistant looks up that index when someone asks. It does not train the model. It requests many URLs, as a search engine does. OpenAI says its search crawler controls appearance in ChatGPT Search answers. Anthropic says its search crawler indexes content for Claude Search.

An AI Agent is a client that works on behalf of a user. It fetches a URL from a question, a paste, or a citation. If that page does not answer the question, it tries another URL. Most sessions still stop after one page.

Provider AI Trainer AI Search AI Agent
OpenAI GPTBot OAI-SearchBot ChatGPT-User
Anthropic ClaudeBot Claude-SearchBot Claude-User
Meta Meta-ExternalAgent Meta-WebIndexer Meta-ExternalFetcher
SpaceXAI — grok-agent grok-agent
Microsoft Bingbot Bingbot —

This table does not list every provider. In our logs, Grok uses one name for search and fetch. Bing’s crawler covers both training and search.

Do AI Agents honor robots.txt?

AI Agents do not treat robots.txt the same way as crawlers.

  • OpenAI says robots.txt rules may not apply when a person asked ChatGPT to fetch.
  • Perplexity says its user-invoked fetch generally ignores the file.
  • Anthropic says Claude honors robots.txt when a person asked.
  • A Chrome-like agent has no bot name, so robots.txt cannot have named directives for it.

If you want AI Trainers off your site, see robots.txt blocks.

AI Trainers and AI Search hit ordinary WordPress sites whether anyone asked about those sites or not. AI Agents fetch URLs when people already want those pages. A GEO file does not attract AI Agents.

Do AI Agents start at your homepage?

No. An AI Agent starts on a specific page, not the homepage.

AI Trainers and AI Search can start at the homepage. They then request more pages. AI Trainers take the content. AI Search adds pages to its index.

An AI Agent looks further when the first page does not answer the question. In the traffic we measured, most sessions still stopped after one URL. The first page was enough.

We prompted various coding assistants the same questions about The SEO Framework. They fetched pricing, the refund policy, a knowledge base article, and our About Us page. They never fetched the homepage. When a guess missed, they hit 404 and tried other slugs. Search didn’t send them to these pages first. They already knew the site from training data. For your site to be in that set of URLs, it still has to be found and indexed. That is SEO.

Landing pages are more important than ever. Improve each page agents tend to hit: the homepage, core product and pricing pages, knowledge base articles, and the About Us page. People paste those URLs. ChatGPT Search and Claude Search cite them. AI Agents guess them.

The llms.txt file

llms.txt is a proposed file at the root of your site. Plugins sell it as an AI ranking lever.

We first tested it in September 2025. We tested it again in September 2026.

Conclusion: llms.txt is not how ChatGPT, Claude, or Grok find or rank your site.

Technology adoption scanners and SEO-tooling crawlers request the file. Those bots check who uses the file, or they collect site data for SEO reports. They do not rank your site in search. In our logs, no AI Agent asked for it. Coding agents can still fetch the file when a person or tool already pointed them at a documentation site. That is not ChatGPT Search ranking your WordPress page. That is a specific user request.

Yoast SEO, Rank Math, AIOSEO, and SEOPress generate llms.txt. Their docs say it helps ChatGPT or “AI” understand the site and get it cited. AIOSEO turns it on by default–their plugin arrives on your site the same way: without consent.

Who requests the llms.txt file?

The public pitch for the file is llmstxt.org. That site is a proposal. Documentation platforms publish the file so tools already pointed at those docs can pull it. That’s not proof that ChatGPT Search fetches /llms.txt on your WordPress site.

Google Search ignores llms.txt. The file does not affect rankings, including AI Overviews and AI Mode.

On the TSF network, we saw about 270 GET /llms.txt requests. On ordinary WordPress sites on the same host, we saw about 300. Those hits were technology adoption scanners, not AI. Zero came from ChatGPT, Claude, or Grok.

This matches GitHub issue 732. After we placed a dummy file, the April 2026 report still showed zero AI-bot hits. Roots.io posted 57 days of logs: about 80,000 AI hits, 101 requests for the file, zero from AI bots.

Independent log studies land in the same place.

  • Ahrefs measured about 38,000 valid llms.txt files across 137,210 domains in May 2026. 97% of those files got zero requests. Of the traffic that reached those files, AI retrieval bots accounted for about 1.1%.
  • Limy counted 408 /llms.txt requests in 515 million AI-bot events.

Should you ship the llms.txt file now?

There is no early-mover advantage. The rush to ship llms.txt now is marketing, not strategy. We will implement it when an AI provider announces they consume it. A PageSpeed Insights note is not that announcement.

When that happens, the provider will fetch the file on your site that day. You do not need it sitting there beforehand.

An AI Trainer will not use llms.txt during training. It already fetches pages from links throughout your site. It trains on that content, not on a curated list of excerpts.

Would the llms.txt file help if AI providers consumed it?

This is a bet, not a measurement. It may be wrong next year.

Even if every assistant started reading llms.txt tomorrow, the file would not move fact-check queries. An llms.txt list of your WordPress pages can help a model describe your site when someone already typed your brand name. It does not win “best X” or “X vs Y.” Those answers get assembled from pages you do not own. The model is not reaching for your llms.txt file. It is pulling the threads and roundups where other people argued about your claims.

Assistants also want a second source. Grok, from our extensive experience, does not treat one commercial page as enough. Court records and government sites are the exception — those get trusted without the same cross-check. Your llms.txt is a self-description of the site. It is not a court record, and it is not a government page. It won’t get cited; it’s just an extra hop.

Why does PageSpeed Insights want the llms.txt file?

PageSpeed Insights only suggests. You do not need the file. PageSpeed Insights uses Lighthouse. A Lighthouse note is not Google Search.

Lighthouse now has an experimental Agentic Browsing category. One audit fetches /llms.txt. Chrome’s docs state that the file is optional. PageSpeed Insights shows that audit.

Google Search still ignores the file. AI Overviews and AI Mode do not use it.

Your WordPress site may still get a red warning from Lighthouse. That note is a lab check. It does not change rankings. It does not make an AI Agent fetch the file. The note may warn that models cannot train on your site. That is Lighthouse copy. It is not a Search ranking document.

Shipping a stub to silence the note can fail the audit. An empty file scores worse than no file.

We will not add a generator so a PageSpeed check turns green.

Use PageSpeed Insights to spot performance issues on your site. A slow site gets bounced more often. Some agents and embeds fail if your site responds too slowly. That is the job. The llms.txt audit is not. See bounce rate reduction.

Why does PageSpeed Insights annotate WebMCP?

WebMCP is a proposed Chrome API. PageSpeed Insights is a Google product. WebMCP registers in-page tools so an agent in the browser can fill a form, book a table, or check out. It is not ranking.

A blog post, a knowledge base article, and most WordPress marketing pages have no tool to register. TSF will not add WebMCP attributes so a PageSpeed check turns green.

Should you serve Markdown to AI clients?

Some AI Agents send the HTTP header Accept: text/markdown. That header asks your site for a Markdown copy of the page instead of HTML. We keep a separate log of those requests on the TSF network (April through September 2026). This table lists only the requests that asked for Markdown.

Who Hits What they’re doing
ShapBot (Parallel) 3,813 Crawler harvest, Markdown preferred.
Claude-User (Claude Code) 1,833 AI Agent, Markdown first.
grok-agent 17 AI Agent, Markdown first.
ChatGPT-User 0 Doesn’t negotiate Markdown in our logs.

There’s a public pitch for Accept: text/markdown at acceptmarkdown.com. From our logs, coding agents send it. ChatGPT’s browse tool didn’t.

Google and Bing do not send Accept: text/markdown. AIOSEO already serves a Markdown copy for “AI engines” at a .md URL. An agent that already fetched the HTML page still has to discover that extra endpoint. That is a second crawl. That spends more tokens, not fewer. SEOPress ships Markdown output too.

Serving Markdown to one client and HTML to the next requires Vary: Accept. That header poisons page caches. A poisoned cache serves Markdown to your visitors. Shortcodes, forms, and paywalls leak into that Markdown.

Agents already extract the page they fetch. They convert HTML on their side. They do not need a Markdown copy from your WordPress site. The converter depends on the request. The agent adjusts. TSF will not add this as a plugin feature.

Does semantic HTML make agents extract the article?

No. Landmarks such as <main> and <article> help browsers and humans parse the page. We assumed AI Agent tools extracted those landmarks. They do not.

The AI Agent tools we tested extracted visible text from URLs we pasted. Even with the <main> and <article> landmarks, header, navigation, breadcrumbs, and the footer still landed in the paste. All of that sits outside both landmarks. Those tools also stripped scripts and styles, including JSON-LD that contains Schema.org. Google Search and Bing still use Schema.org, but AI Agents do not.

This is our prophecy: Agents ought, should, and will build better tooling around the semantic web. The Semantic Web is the idea that a page should mark what each part means, so a machine can use it. HTML landmarks and Schema.org are that contract today. You should already put the complete answer to the page’s topic in <article>. Well-written WordPress themes do that for you. It helps with accessibility; screen readers use those landmarks. Google Search asks you to make the main content easy to distinguish from the rest of the page. Then a snippet can come from the article, not the menu.

Elementor, as the most popular page builder, deserves a callout: We found that self-hosted Elementor pages are not semantically valid HTML. A post sits as one widget among other widgets. That layout can omit <main> and <article>. Our prophetic converter that wants “the article” won’t be able to find it.

Yet llms.txt keeps winning the debate: It’s easy to implement and forget about. The cost of the file is not that it is useless. The cost is what it displaces. An hour on the file is an hour not spent improving your pages. Writing the page takes more time. It’s harder to show as progress. That’s why speculative tricks keep getting shipped.

What should you do instead?

  • Ignore AI flags on checklist tools. PageSpeed Insights, Lighthouse, and SEO plugins will highlight llms.txt and GEO scores. Those checks are bogus.
  • Write the landing pages. The homepage, core product pages, knowledge base articles, About Us, and pricing get the fetches. Put the complete answer on each. An AI Agent arrives with a URL from search, a paste, a citation, or a guessed WordPress-ish slug. The page it fetches is the document. See practical SEO tips.
  • One canonical URL. Search follows the canonical. When an AI Agent arrives from search, that is the URL it fetches. It won’t fetch the Markdown URL. Do not invent a second URL for “AI.” We have honed canonical URLs for over a decade. For example, see advanced query protection.
  • Use semantic HTML. Put the article in <main> / <article>. Mark header, nav, and footer with data-nosnippet if they should stay out of Google Search snippets.
  • Structured data. Keep Schema.org for Google Search. The extractors we tested stripped scripts and styles, including JSON-LD. That is not a GEO input. See structured data supported by The SEO Framework.
  • robots.txt. The AI crawler block in SEO Settings → Robots Settings → Robots.txt is optional. Use it if you want AI Trainers off the site. That block does not stop ChatGPT Search or ChatGPT’s fetcher. See robots.txt blocks.

Will The SEO Framework add AI SEO features?

Don’t count on it.

Feature Why not
GEO score or AEO checklist Those labels are marketing names. On Google Search, they still mean traditional SEO.
llms.txt AI Agents don’t request it. Google Search ignores it. We’ll add it when a provider announces they consume it.
Markdown for AI clients From our logs, ChatGPT’s browse tool didn’t ask. Serving both formats poisons page caches. AI Agents already extract the page they fetch; they convert HTML on their side.
WebMCP or a PageSpeed Agentic score A PageSpeed check is not a ranking factor.

An “SEO score” teaches you how to chase the score, not how to write the page. We skipped readability scores for the same reason.

TSF still ships robots.txt, canonicals, redirects, structured data, and tons of other features that are actually helpful. TSF also generates titles and descriptions without needing expensive and intrusive AI features.

What might change later?

This section is speculation. It may be wrong next week.

An AI provider may one day announce they consume llms.txt. That announcement is the signal. A plugin generator, a Google AI-readiness page, and a Lighthouse note are not. Consumption still would not make the file rank your website better.

Until then, write the complete answer on the public WordPress page. An AI Agent fetches that URL. If the page answers the query, the agent will use that page.

Filed Under: SEO, Robots, Structured Data

Related articles

  • SEO

    • Practical SEO tips
  • Robots

    • Advanced Query Protection
    • Robots.txt blocks
    • Why aren’t my empty categories indexed?
  • Structured Data

    • WooCommerce integration
    • Breadcrumb shortcode
    • Structured data supported by The SEO Framework

Commercial

The SEO Framework
Trademark of CyberWire B.V.
Leidse Schouw 2
2408 AE Alphen a/d Rijn
The Netherlands
KvK: 83230076
BTW/VAT: NL862781322B01

Twitter  GitHub

Professional

Pricing
About
Support
Terms and Conditions
Refund Policy

Editorial

Knowledge Base
Release Notes
Feature Highlights
Blog
Privacy Policy

Practical

TSF on WordPress
TSF on GitHub
TSFEM on here
TSFEM on GitHub
Deploy Troy

Happy customers in 2026 › The SEO Framework