How to Block Aggressive AI Crawlers Like Bytespider on WordPress
Quick Answer: Bytespider uses over 100,000 IP addresses and frequently ignores robots.txt, making traditional blocking methods ineffective for WordPress sites. Server-level user-agent blocking in .htaccess or nginx catches identified crawlers before WordPress loads, while Cloudflare’s edge-level AI Training block enforces restrictions at the connection level. WordPress plugins like citelayer track 62 distinct AI bot signatures. The key distinction is blocking training crawlers while keeping search and citation bots that drive referral traffic.
In this article
- Why does robots.txt fail against aggressive AI crawlers?
- Which AI crawlers actually respect robots.txt?
- How do you block Bytespider on WordPress without a CDN?
- How does Cloudflare’s AI Training block help against aggressive crawlers?
- What server resources do aggressive AI crawlers consume on WordPress?
- Should you block all Chinese AI crawlers?
- What is the complete robots.txt disallow list for AI training crawlers?
- How do you monitor which AI crawlers are hitting your WordPress site?
Why does robots.txt fail against aggressive AI crawlers?
Robots.txt is a voluntary protocol — compliant crawlers check it, but Bytespider and several Chinese AI crawlers frequently ignore it, consuming server resources and scraping content regardless of your directives. The protocol assumes good faith. It tells a crawler what not to do, but it has zero enforcement mechanism. A crawler that decides to ignore the file faces no technical barrier.
Here’s the thing… Bytespider uses over 100,000 IP addresses, making IP-level blocking impractical without a CDN or WAF (Chinese AI Blocking Guide, 2026). Even if you tried blocking IPs one by one, the volume makes it a losing game. The addresses rotate, new ranges appear, and your firewall rules become unmanageable within days. Robots.txt was designed for an era when crawlers were fewer, slower, and polite. That era ended when AI training demand turned web scraping into an industrial operation.
Bytespider uses over 100,000 IP addresses, making IP-level blocking impractical without a CDN or WAF (Chinese AI Blocking Guide, 2026).
Which AI crawlers actually respect robots.txt?
GPTBot, ClaudeBot, Google-Extended, and AppleBot-Extended all respect robots.txt directives — but Bytespider, some instances of DeepSeekBot, and several smaller Chinese crawlers frequently do not. The difference is transparency and accountability. OpenAI publishes IP ranges and respects robots.txt; Bytespider publishes neither reliably (Chinese AI Blocking Guide, 2026).
This creates a two-tier system. Western AI companies — OpenAI, Anthropic, Google, Apple — built their crawlers to honour the protocol because they operate under regulatory scrutiny and reputational pressure. ByteDance’s Bytespider, by contrast, operates with less transparency about its crawling practices. The result is that a robots.txt file written only for the compliant tier leaves your content wide open to the non-compliant one.
| Crawler | Respects robots.txt | Publishes IP ranges | Purpose |
|---|---|---|---|
| GPTBot (OpenAI) | Yes | Yes | Training + search |
| ClaudeBot (Anthropic) | Yes | Yes | Training |
| Google-Extended | Yes | Yes | Training |
| Bytespider (ByteDance) | Unreliable | No | Training |
| CCBot (Common Crawl) | Yes | Partial | Training dataset |
| DeepSeekBot | Unreliable | No | Training |
Related: WordPress Tracking articles
How do you block Bytespider on WordPress without a CDN?
Add user-agent blocking rules in your .htaccess file or nginx configuration to reject requests from Bytespider, CCBot, and other known aggressive crawlers at the server level before they consume PHP resources. Server-level UA blocking in .htaccess or nginx catches identified crawlers before WordPress even loads — zero PHP overhead (Chinese AI Blocking Guide, 2026).
For Apache (.htaccess), the rule is straightforward. You match the user-agent string and return a 403 Forbidden before WordPress or any PHP process is even invoked. The request dies at the web server layer. For nginx, you add a condition in the server block that checks the $http_user_agent variable and returns 403 for matching strings.
Translation: your WordPress installation never knows the crawler was there. No PHP execution cycle, no database queries, no plugin hooks fired. The web server handles the rejection itself, which is why this approach costs essentially nothing in performance terms — even at high request volumes.
The limitation is that user-agent strings can be spoofed. A crawler that lies about its identity will pass straight through UA-based blocking. This is where layered defences matter — server-level UA blocking catches the honest-but-aggressive crawlers, and a WAF or CDN catches the rest.
How does Cloudflare’s AI Training block help against aggressive crawlers?
Cloudflare’s Training category block operates at the network edge before requests reach your server — it catches crawlers that ignore robots.txt because the block is enforced at the connection level, not as a voluntary protocol. Cloudflare processes approximately 20% of all web traffic, making its edge-level blocking the most effective anti-crawler tool available to WordPress sites (Help Net Security / Cloudflare, 2026).
In July 2026, Cloudflare introduced three-category AI crawler controls: Search (bots that drive referral traffic), Training (bots that scrape for model training), and Fetching (bots that retrieve content for real-time AI answers). The granularity matters. You can block Training crawlers — including Bytespider — while keeping Search and Fetching bots that actually send visitors to your site.
Cloudflare processes approximately 20% of all web traffic, making its edge-level blocking the most effective anti-crawler tool available to WordPress sites (Help Net Security / Cloudflare, 2026).
The key advantage over robots.txt is enforcement. Robots.txt asks a crawler to comply. Cloudflare blocks the connection before the request ever reaches your origin server. A crawler that ignores your robots.txt cannot ignore a TCP connection that Cloudflare refuses to open.
What server resources do aggressive AI crawlers consume on WordPress?
Each crawler request triggers a full WordPress PHP execution cycle — database queries, theme rendering, plugin hooks — and at Bytespider’s crawl rates across 100,000+ IPs, the load can slow your store for real customers or trigger hosting overage charges. A single aggressive crawler can generate more requests per day than your entire human traffic — consuming hosting resources you are paying for (Chinese AI Blocking Guide, 2026).
WordPress is not a static file server. Every page request that reaches PHP triggers the full application stack: database connection, query execution, theme template processing, and every active plugin’s hooks. On a WooCommerce store, add product data queries, cart state checks, and pricing calculations to that list. Multiply by thousands of requests per hour from an aggressive crawler, and you have a denial-of-service scenario where the attacker is a bot that thinks it’s being helpful.
The cost is direct. Shared hosting plans throttle or suspend accounts that exceed CPU limits. Managed WordPress hosts charge for pageview overages. Even on a VPS or dedicated server, the CPU and memory consumed by crawler traffic is capacity that could be serving paying customers. Let that sink in.
Related: More WordPress Tracking guides
Should you block all Chinese AI crawlers?
Block training crawlers (Bytespider, CCBot) but keep search/citation bots (Baiduspider in search mode) if you sell to Chinese-speaking markets — the same Search vs Training distinction Cloudflare now enforces applies to your robots.txt strategy. Alibaba’s Qwen has 113,000+ derivative models on HuggingFace — blocking Qwenbot blocks potential agent commerce from the Chinese AI ecosystem (Chinese AI Blocking Guide / HuggingFace, 2026).
The question isn’t whether to block Chinese crawlers. The question is which ones. A blanket block is simple but commercially reckless if you have any Chinese-speaking audience. Baiduspider in search mode drives actual referral traffic from Baidu search results. Qwenbot could eventually drive agent commerce from Alibaba’s AI ecosystem. Blocking them is blocking potential revenue.
The strategic approach mirrors what Cloudflare codified: separate Training from Search. Block Bytespider, ChatGLM-Spider, PanguBot, and any crawler whose sole purpose is scraping content for model training. Keep Baiduspider (search mode), Qwenbot (if agent commerce matters to you), and any bot that sends real visitors.
What is the complete robots.txt disallow list for AI training crawlers?
A comprehensive robots.txt block for training crawlers should disallow GPTBot, CCBot, Bytespider, ChatGLM-Spider, PanguBot, Google-Extended, FacebookBot, and Omgilibot — while explicitly allowing OAI-SearchBot, PerplexityBot, and ClaudeBot which drive referral traffic. 28+ known AI crawler user-agent strings are documented in the GeoPromptTracker dataset — your robots.txt should address each one individually (Chinese AI Blocking Guide, 2026).
The mistake most WordPress site owners make is writing a partial list and assuming it covers everything. The AI crawler landscape in 2026 has over 28 documented user-agent strings, and new ones appear regularly. A robots.txt that only blocks GPTBot and CCBot leaves dozens of training crawlers with unrestricted access.
Equally important: do not block everything. OAI-SearchBot (OpenAI’s search crawler, distinct from GPTBot), PerplexityBot, and ClaudeBot’s citation functions drive real referral traffic. Blocking them removes your content from AI-powered search results — the opposite of what an AEO strategy needs. The robots.txt file is not a firewall; it is a policy document. Use it to express which crawlers you welcome and which you do not, knowing that the non-compliant ones need server-level or CDN-level enforcement anyway.
How do you monitor which AI crawlers are hitting your WordPress site?
Install the Known Agents (Dark Visitors) plugin or citelayer to surface AI crawler activity in your WordPress dashboard — then cross-reference with raw server access logs to catch crawlers that slip past plugin detection. citelayer tracks 62 distinct AI/LLM bot signatures; Known Agents integrates with the Dark Visitors database for broader coverage (WordPress.org, 2026).
Plugin-level monitoring has a structural limitation: it only sees requests that reach WordPress. A crawler blocked at the server level (via .htaccess or nginx rules) or at the CDN edge (via Cloudflare) never triggers a WordPress hook, so the plugin never knows it existed. That is a good thing for blocking, but it means your monitoring picture is incomplete unless you also check raw server logs.
The practical workflow: install citelayer or Known Agents for dashboard-level visibility into which AI bots reach your WordPress application layer. Then periodically review your raw access logs (or a log analysis tool like GoAccess) for user-agent strings that match known AI crawlers but bypassed your plugin detection. The combination gives you both real-time dashboard visibility and a ground-truth audit trail.
Key Takeaways
- Robots.txt is voluntary, not enforced: aggressive crawlers like Bytespider ignore it, so server-level and CDN-level blocking are essential.
- 100,000+ IPs make IP blocking futile: user-agent blocking at the server or edge level is the practical alternative.
- Cloudflare’s three-category model is the standard: block Training crawlers, keep Search and Fetching bots that drive traffic.
- Server-level UA blocking costs zero PHP overhead: .htaccess or nginx rules reject crawlers before WordPress loads.
- 28+ AI crawler user-agents exist in 2026: a partial robots.txt list leaves your site exposed.
- Block training, keep search: do not blanket-block all AI crawlers if you want AI-powered search visibility.
- Monitor with plugins and raw logs together: citelayer tracks 62 bot signatures, but only at the WordPress layer.
Robots.txt is a voluntary protocol — compliant crawlers check it, but Bytespider and several Chinese AI crawlers frequently ignore it, consuming server resources and scraping content regardless of your directives.
GPTBot, ClaudeBot, Google-Extended, and AppleBot-Extended all respect robots.txt directives — but Bytespider, some instances of DeepSeekBot, and several smaller Chinese crawlers frequently do not.
Add user-agent blocking rules in your .htaccess file or nginx configuration to reject requests from Bytespider, CCBot, and other known aggressive crawlers at the server level before they consume PHP resources.
Cloudflare’s Training category block operates at the network edge before requests reach your server — it catches crawlers that ignore robots.txt because the block is enforced at the connection level, not as a voluntary protocol.
Each crawler request triggers a full WordPress PHP execution cycle — database queries, theme rendering, plugin hooks — and at Bytespider’s crawl rates across 100,000+ IPs, the load can slow your store for real customers or trigger hosting overage charges.
Block training crawlers (Bytespider, CCBot) but keep search/citation bots (Baiduspider in search mode) if you sell to Chinese-speaking markets — the same Search vs Training distinction Cloudflare now enforces applies to your robots.txt strategy.
A comprehensive robots.txt block for training crawlers should disallow GPTBot, CCBot, Bytespider, ChatGLM-Spider, PanguBot, Google-Extended, FacebookBot, and Omgilibot — while explicitly allowing OAI-SearchBot, PerplexityBot, and ClaudeBot which drive referral traffic.
Install the Known Agents (Dark Visitors) plugin or citelayer to surface AI crawler activity in your WordPress dashboard — then cross-reference with raw server access logs to catch crawlers that slip past plugin detection.