← Back to Blog

robots.txt for AI in 2026: What to Block, What to Allow, and Why the Old…

Quick Answer: The old approach of blocking or allowing ‘AI bots’ as a single category is wrong because it treats three fundamentally different visitors the same — training crawlers that take content, search bots th (HUMAN Security, 2026). Allow search and citation bots — OAI-SearchBot, PerplexityBot, ChatGPT-User, GoogleOther, and ClaudeBot — because they drive referral traffic and prod. This article covers the full picture — from why this matters to what to do about it.

Why does the old robots.txt strategy not work for AI bots?

The old approach of blocking or allowing ‘AI bots’ as a single category is wrong because it treats three fundamentally different visitors the same — training crawlers that take content, search bots that send traffic, and shopping agents that send revenue. Training crawlers accounted for 74% of AI traffic but generate zero revenue; search bots and agents generate traffic and purchases (HUMAN Security, 2026).

This isn’t a marginal difference or a measurement artifact. It’s a structural gap in how tracking works — and it affects every metric downstream, from ROAS to bidding model quality. The fix isn’t incremental; it requires a different architectural approach.

Training crawlers accounted for 74% of AI traffic but generate zero revenue; search bots and agents generate traffic and purchases (HUMAN Security, 2026).

Which AI bots should WooCommerce stores allow in robots.txt?

Allow search and citation bots — OAI-SearchBot, PerplexityBot, ChatGPT-User, GoogleOther, and ClaudeBot — because they drive referral traffic and product visibility in AI-powered search results. ChatGPT commands 76.85% of AI referral share; blocking OAI-SearchBot removes you from that channel (AuthorityTech / Statcounter, 2026).

Understanding where the separation happens is the first step toward building a reliable measurement system. The architecture dictates the options — and most of the options that seem obvious at first turn out to be dead ends once you look at where the data actually flows.

Which AI bots should WooCommerce stores block in robots.txt?

Block training-only crawlers — GPTBot (training mode), CCBot, Bytespider, ChatGLM-Spider, PanguBot, Google-Extended, FacebookBot, and Omgilibot — which scrape content for model training and return nothing. 28+ known AI crawler user-agent strings documented; each needs individual robots.txt treatment (Chinese AI Blocking Guide, 2026).

The routing logic is straightforward once the tagging is in place. The complexity isn’t in the routing — it’s in getting the tag right at the point of origin, before the data enters any downstream system.

28+ known AI crawler user-agent strings documented; each needs individual robots.txt treatment (Chinese AI Blocking Guide, 2026).

Related: Google’s Qualified Future Conversions Metric Counts Sales Up to 180 Days After a Click — What WooCommerce Store Owners Need to Change in Their Tracking Setup

What is the difference between GPTBot and OAI-SearchBot?

GPTBot scrapes content for training OpenAI’s models (block it); OAI-SearchBot fetches pages in real-time to answer user queries in ChatGPT search (allow it — it sends you traffic). OpenAI operates at least 4 distinct user-agents: GPTBot, OAI-SearchBot, ChatGPT-User, and ChatGPT Agent — each with a different purpose (HUMAN Security, 2026).

The cascade effect is what makes this urgent rather than merely important. Each day of mixed data compounds the problem — the algorithm trains on contaminated signals, and undoing that training takes longer than preventing it.

Does robots.txt apply to AI shopping agents?

No — browser-based shopping agents run real Chromium engines and send standard Chrome user-agent strings, so they do not check robots.txt at all and cannot be controlled through it. 71% of agentic activity is browser-based — they bypass robots.txt entirely by design (AgentLux / HUMAN Security, 2026).

The distinction between detection and control is critical. Many store owners install a detection plugin and assume the problem is solved. But detection without routing control is observation without action — you know what happened, but you haven’t changed what your ad platforms learn from it.

How do you handle DeepSeek in robots.txt when it has no user-agent?

You cannot — DeepSeek does not publish a crawler user-agent string, so there is nothing to disallow in robots.txt. The only options are IP range blocking at the server level or Cloudflare’s AI controls. DeepSeek is the only major AI vendor with no published crawler identity — it crawls invisibly (xSeek, 2026).

This is a common misconception — that adding a layer between WordPress and the ad platforms solves the tagging problem. It doesn’t, because the layer operates on data it receives, and if that data arrives untagged, no amount of downstream processing can recover the missing label.

Related: Five GA4 Volume Thresholds Your WooCommerce Store Fails — And Each One Makes the Others Worse

What does a proper 2026 robots.txt for WooCommerce look like?

A 2026 robots.txt has three sections: explicit Allow directives for search bots you want traffic from, explicit Disallow directives for training crawlers, and no mention of shopping agents because they do not read robots.txt. The optimal configuration blocks training (zero revenue) while preserving search visibility (referral traffic) and ignoring agents (high-converting revenue) (Chinese AI Blocking Guide, 2026).

The good news is that the setup is bounded. You don’t need to rebuild your entire tracking stack. You need three specific components, each with a clear job, and they layer on top of whatever you’re already running.

Related: Your WooCommerce Consent Banner Rejection Rate Is 40-70% in the EU — That’s Not Lost Traffic, It’s Unmeasured Revenue

30 DAY FREE TRIAL

No card needed. Take a strong step to getting into Data Heaven today!

Let's Do It !

How often should you update your AI robots.txt rules?

Review quarterly — new AI crawlers emerge monthly, naming conventions change, and previously compliant bots sometimes stop respecting directives, so a robots.txt written in January 2026 is already missing agents that launched since. At least 6 new agent-related WordPress plugins launched in the first half of 2026 alone — the ecosystem is moving faster than annual reviews can track (WordPress.org, 2026).

This is where the investment pays off in operational clarity. Instead of one confused number driving one confused decision, you get two clean signals driving two distinct strategies — each optimised for the audience it actually serves.

LayerWhat it doesLimitation
Client-side pixelFires in the browser on page eventsBlind to API/MCP purchases — no browser session exists
Detection pluginTags orders as agent or human in WordPressDoesn’t control what ad platforms receive
Server-side hookInspects origin at order creation, writes cohort flagRequires implementation at the PHP level
Server-Side GTMRoutes events to endpoints conditionallyCan’t tag events — only route what’s already tagged

FREE 30 DAY TRIAL

Take a strong step to getting into Data Heaven today! No card needed.

Start NOW !

Key Takeaways

  • Why does the old robots.txt strategy not work for AI bots: The old approach of blocking or allowing ‘AI bots’ as a single category is wrong because it treats.
  • Which AI bots should WooCommerce stores allow in robots.txt: Allow search and citation bots — OAI-SearchBot, PerplexityBot, ChatGPT-User, GoogleOther, and.
  • Which AI bots should WooCommerce stores block in robots.txt: Block training-only crawlers — GPTBot (training mode), CCBot, Bytespider, ChatGLM-Spider, PanguBot,.
  • What is the difference between GPTBot and OAI-SearchBot: GPTBot scrapes content for training OpenAI’s models (block it); OAI-SearchBot fetches pages in.
  • Does robots.txt apply to AI shopping agents: No — browser-based shopping agents run real Chromium engines and send standard Chrome user-agent.
  • How do you handle DeepSeek in robots.txt when it has no user-agent: You cannot — DeepSeek does not publish a crawler user-agent string, so there is nothing to disallow.
Why does the old robots.txt strategy not work for AI bots?

The old approach of blocking or allowing ‘AI bots’ as a single category is wrong because it treats three fundamentally different visitors the same — training crawlers that take content, search bots that send traffic, and shopping agents that send revenue.

Which AI bots should WooCommerce stores allow in robots.txt?

Allow search and citation bots — OAI-SearchBot, PerplexityBot, ChatGPT-User, GoogleOther, and ClaudeBot — because they drive referral traffic and product visibility in AI-powered search results.

Which AI bots should WooCommerce stores block in robots.txt?

Block training-only crawlers — GPTBot (training mode), CCBot, Bytespider, ChatGLM-Spider, PanguBot, Google-Extended, FacebookBot, and Omgilibot — which scrape content for model training and return nothing.

What is the difference between GPTBot and OAI-SearchBot?

GPTBot scrapes content for training OpenAI’s models (block it); OAI-SearchBot fetches pages in real-time to answer user queries in ChatGPT search (allow it — it sends you traffic).

Does robots.txt apply to AI shopping agents?

No — browser-based shopping agents run real Chromium engines and send standard Chrome user-agent strings, so they do not check robots.txt at all and cannot be controlled through it.

How do you handle DeepSeek in robots.txt when it has no user-agent?

You cannot — DeepSeek does not publish a crawler user-agent string, so there is nothing to disallow in robots.txt. The only options are IP range blocking at the server level or Cloudflare’s AI controls.

What does a proper 2026 robots.txt for WooCommerce look like?

A 2026 robots.txt has three sections: explicit Allow directives for search bots you want traffic from, explicit Disallow directives for training crawlers, and no mention of shopping agents because they do not read robots.txt.

How often should you update your AI robots.txt rules?

Review quarterly — new AI crawlers emerge monthly, naming conventions change, and previously compliant bots sometimes stop respecting directives, so a robots.txt written in January 2026 is already missing agents that launched since.

References

  1. HUMAN Security (2026). Source material. Source
  2. AuthorityTech / Statcounter (2026). Source material. Source
  3. Chinese AI Blocking Guide (2026). Source material. Source
  4. AgentLux / HUMAN Security (2026). Source material. Source
  5. xSeek (2026). Source material. Source
  6. WordPress.org (2026). Source material. Source