The Spectrum Dispatch News

technology

Website Owner Battles Bot Traffic: 99% of Requests Are Automated

A developer shares tactics for fighting scrapers and AI crawlers on a 1.5 million-page philanthropy database, from country blocks to rate limits.

Website Owner Battles Bot Traffic: 99% of Requests Are Automated

A web developer running PatronView, a 1.5 million-page database of American philanthropists, has spent a year combating bot traffic on the site. According to the author, server logs reveal a stark discrepancy: while visitor analytics tools recorded about 5,977 pageviews in one week, the server processed 2.5 million requests total, with roughly 214 bot page loads for every human-viewed page.

Website Owner Battles Bot Traffic: 99% of Requests Are Automated

The bot problem escalated dramatically in April 2025, when the site received 3.6 million requests in a single day from 361,844 unique IP addresses, primarily originating from China. Cloudflare’s Managed Challenge blocked 1.18 million of those requests within ten hours, but many bots still bypassed the challenge. The author responded by blocking traffic from China, Vietnam, and Singapore entirely at the edge—countries that sent no meaningful legitimate traffic to an English-language American database.

Separate from geographic attacks, the site faced persistent scraping from AI training crawlers. According to the author’s measurements, Anthropic’s Claude-SearchBot maintained a ratio of 35,000 pages crawled per human visitor referred, compared to Cloudflare’s reported 3,000-to-1 ratio for Anthropic crawlers generally. In June, Claude-SearchBot requested 420,680 pages while sending only 12 human visitors, yet consumed 4.63 GB of bandwidth against 175 KB for the traffic it generated. When the author blocked the crawler with a 403 (access denied) response, requests dropped from 60,000 daily to 25 attempts per day, suggesting the crawler respects such signals.

Amazon’s Amzn-SearchBot proved even more aggressive, reaching 117,000 requests daily without providing meaningful referral traffic. The author blocked it after discovering it feeds Rufus and Alexa results without sending visitors or attribution.

In July, the site faced attacks from headless Chrome browsers running on AWS datacenters. The author implemented firewall rules challenging all datacenter traffic, since legitimate users browse from residential ISPs like Comcast and T-Mobile, not cloud infrastructure. The author also discovered that Cloudflare’s JavaScript Detection feature was costing 2,875 milliseconds of page load time on mobile devices, reducing Lighthouse scores by 40 points while providing unreadable telemetry. Disabling it improved mobile scores from 58 to 99.

The author acknowledges the irony of fighting scrapers while using web scraping to build PatronView’s own database from public IRS forms and donor walls. A survey of other developers on social media revealed widespread similar struggles, with country-blocking becoming common practice and some developers abandoning bot-fighting efforts due to complexity.

Key facts

  • Server logs showed 2.5 million total requests per week versus 5,977 recorded pageviews—a 214:1 bot-to-human ratio
  • 3.6 million requests arrived in a single day from 361,844 Chinese IP addresses in April 2025
  • Claude-SearchBot crawled 35,000 pages per human visitor referred, consuming 4.63 GB bandwidth for minimal traffic
  • Amazon’s Amzn-SearchBot reached 117,000 requests daily without sending web traffic or referrals
  • Blocking datacenter IPs reduced attacks from headless browser fleets on AWS infrastructure
  • Cloudflare’s JavaScript Detection feature cost 2.875 seconds of mobile page load time

Sources

← All posts