A developer managing PatronView, a 1.5 million-page database of American philanthropists, has spent a year battling automated scrapers and bots. According to his detailed account, the scale of the problem is staggering: server logs show 214 bot page loads for every 1 human page load, despite visitor analytics tools showing only 5,977 pageviews in a week when the server actually answered 2.5 million requests.

The attacks escalated throughout the year. In November 2025, a coordinated botnet of 4,000 scrapers appeared, each hitting only a single page type. Then in April, the site received 3.6 million requests in a single day from 361,844 unique IP addresses, nearly all originating from China. Cloudflare’s CAPTCHA challenge absorbed 1.18 million of those requests in the first ten hours, yet an “alarming number” still passed the verification.
Faced with the volume, the developer blocked China, Vietnam, and Singapore at the network edge—countries that represent only 0% of his intended audience. The move proved necessary after what he describes as a “copycat wave” from Vietnam began immediately after the initial Chinese attack.
AI-powered crawlers presented a different problem. Claude-SearchBot crawled 420,680 pages in one week while sending only 12 human visitors, yielding a ratio of 35,000 pages crawled per visitor referred. Amazon’s AI search crawler, Amzn-SearchBot, escalated to 117,000 requests per day with zero referrals. Both were blocked with simple firewall rules, and according to the account, these polite AI crawlers respected the 403 access denial response immediately.
Other aggressive crawlers included SEO tools like SemrushBot, AhrefsBot, and MJ12bot, which were blocked at the Cloudflare level. Bing, by contrast, was permitted because it delivered growing referral traffic despite crawling at 158,610 requests for 680 visitors.
The developer’s most effective strategy involved blocking all datacenter IP addresses—AWS, Azure, and similar cloud infrastructure—since legitimate human readers browse from residential ISPs like Comcast and T-Mobile. This rule blocked headless Chrome browsers being run at scale.
Optimization attempts had unintended consequences. A JavaScript detection script from Cloudflare, intended to challenge bots, added 2,875 milliseconds of latency on mobile devices and reduced the site’s Lighthouse performance score to 58. Removing it improved the score to 99 within an hour, though it also removed a layer of bot defense.
The CAPTCHA solve rate on the site measured only 0.24%, suggesting bots rarely attempt to solve challenges and instead abandon requests. The developer notes the irony that his site scrapes public documents itself while fighting scrapers targeting his aggregated data.
Key facts
- Server received 2.5 million requests in one week; visitor stats recorded only 5,977 pageviews
- 3.6 million requests arrived in a single day from Chinese IP addresses in April
- Claude-SearchBot crawled 35,000 pages per visitor referred; Amazon’s crawler reached 117,000 requests per day
- CAPTCHA solve rate was only 0.24%, indicating bots rarely attempt verification
- Blocking entire countries (China, Vietnam, Singapore) was necessary due to volume and lack of legitimate audience
