The AI Scraper War Escalates: Residential Proxy Networks Are Eating the Open Web
AI scrapers are using millions of residential proxies to pillage web sites. LWN reports on the arms race, the takedowns, and why proof-of-work won't save us.
LWN.net has published a detailed update on the ongoing war between web site operators and AI scraper bots. The piece, written by Jonathan Corbet, paints a grim picture: the scraping problem hasn't abated — it's gotten worse. What started as a nuisance in early 2025 has evolved into a full-scale assault on the open web, powered by residential proxy networks that make traditional IP blocking nearly useless.
How Residential Proxies Work
Scraper attacks now originate from millions of unique IP addresses, each hitting a site only a few times before disappearing. These addresses come from residential and mobile networks, controlled by central command nodes. Software installed on ordinary devices — often without the owner's knowledge — fetches pages on demand and forwards the data back to the controller. The term for this is "residential proxies."
There are two main types of operators running these networks. The first is purely criminal: malware compromises devices, including media-streaming boxes, to build botnets. Google took down one such network, IPIDEA, earlier this year, which briefly reduced scraper traffic at LWN. More recently, a botnet called NetNut was disrupted in coordination with the FBI.
The second type is more insidious: companies like Bright Data offer "ethically sourced" residential proxies. They provide free VPN services or SDKs that app developers can embed, effectively turning users' devices into scraping endpoints. These operators claim GDPR compliance, but the result is the same — your phone becomes a weapon against web sites you may never visit.
Why Blocking Doesn't Work
Bots don't fetch images or CSS, so they can be identified — but by the time you do, the IP is gone. Blocking is futile. The big frontier-model companies (OpenAI, Google, etc.) do their own scraping with identifiable user-agent strings and respect robots.txt, but they still hammer sites repeatedly. The real damage comes from anonymous proxy traffic, whose customers remain unknown.
LWN has implemented aggressive defenses but won't detail them, for obvious reasons. They've avoided proof-of-work tools like Anubis because scrapers can offload the work to their proxy armies. They've also avoided allowlisting dominant search engines, which would only entrench monopolies.
The Arms Race Continues
Every takedown brings a temporary lull, but new networks appear. Google's Play Store now checks for NetNut-infected apps, but the bigger question — why app stores make it so easy to distribute proxyware — remains unanswered. Until the industry behind LLMs is held to a minimal ethical standard, site operators will keep fighting a defensive war they can't win.
Source: LWN.net
Discussion
0 Comments
Be the first to start the discussion.