Concerns about the evidence for the AI scraper problem and the resulting blocking

Xe Iaso <[email protected]>
Newsgroups gmane.comp.gcc.devel
Message-ID <[email protected]>
Hello,

I'm Xe Iaso, the creator of the Anubis project. As such, note my bias 
when I am talking about issues related to this flavour of problems.

I have been working with the Sourceware team for over a year to try and 
minimize impact on users and help notify them of emergent threats as 
they pop up. This includes a lot of behind the scenes advocacy that is 
not generally visible to the public, such as the (futile) work I've been 
doing to try to get the Lightpanda team to respect RFC 9309 among other 
internet standards.

As an aside, while I was reading your email, I noticed signs that it was 
in part generated using a generative AI tool. Charitably, I assume you 
did this to translate your original message from your native language so 
that your email looks "professional". In the future, if you are going to 
use AI to translate emails, please mark it as so and attach the original 
version in your native language.

On 8/24/26 9:16 AM, Kalamatee wrote:
> I have some concerns about the way the current “AI scraper” problem is
> being characterised, and particularly about the measures being taken in
> response to it.

I don't like that this has to be the option either. I have tried many 
things to reduce the level of paranoia involved, but at some level this 
is becoming a big data problem and I really don't want to receive 
additional threats in order to collect that data to make a better 
option. I'm working on a reputation database that I will talk about more 
in the future.

> I don't doubt that Sourceware has seen unusual or excessive traffic, and I
> don't think it is unreasonable to investigate whether automated scraping is
> contributing to load. What I am less convinced about is the extent to which
> the public evidence actually supports the conclusions being drawn from that
> traffic.

As a rule of thumb: if multiple organizations are all coming to the same 
conclusions without colluding, it's probably because a thing is actually 
happening. I have no way to prove this in a way you'll likely accept, 
but I honestly don't have the time or energy to run a grand conspiracy 
like you're imagining.

I would like to know what your hypothesis is as to the source of this 
traffic.

> In particular, there seems to be very little evidence beyond 
> assertions that the traffic being classified as AI scraping is
> actually illegitimate traffic, rather than simply automated traffic. I
> haven't seen request statistics, representative logs, request
> patterns, resource-consumption figures, or anything else that would
> allow people outside the infrastructure team to understand what was
> actually happening and how significant the problem was.

Pedantically, yes, there is no way to prove that these scrapers are for 
getting _generative AI training data_ in particular. The nature of a 
bunch of random Googles Chrome coming from random residential IP ranges 
globally means it's hard to track the threat actors down.

However based on the few scrapers that are honestly identifying 
themselves (eg: the x.ai scraper that seemed precision-designed to make 
GitLab/Gitea/Forgejo servers fall over, or the AmazonBot run by the 
Alexa team) this seems to be a non-unreasonable conclusion.

In this era of internet, everything has to be C10K ready[3] and a lot of 
critical infrastructure just isn't. This isn't something that can be 
easily solved overnight. This is a big enough problem that GitHub and 
other key services are falling over from the load too. Why do you think 
everything is becoming login only? Hint: it's not for fun.

The only real assumption that a lot of us are left with is that it's 
being done by AI companies in an effort to strip the Internet for parts 
to gather as much training data as possible in order to brute force 
improvements.

Speaking as someone that genuinely does have domain expertise in AI 
usage and training in general, this is one of the few methods of getting 
foundation model training data left. The others include scan -> shred 
pipelines to ingest books picked for pruning from libraries and unsold 
inventory from bookstores.

Stripping the Internet for parts is the only real thing left because it 
requires the least investment and capitalism optimizes for low cost.

> The same applies to the suggestion that AI scraping companies are obtaining
> access to large numbers of residential IP addresses through libraries or
> services which effectively sell users' bandwidth. That may well be what is
> happening, but the discussion on the mailing lists appears to present this
> as a hypothesis based on an external lead rather than something
> demonstrated from Sourceware's own traffic data. I don't think we should
> treat that explanation as established fact without evidence.

Speaking as someone with half their foot in the AI industry, people 
actually do this. There are tutorials spread with paid notion notebook 
shares that go into explicit detail on how to do this. When the data[1] 
shows 2.6 million IP addresses covering 229 countries (for context: ISO 
3166-1 comprises 249 geographic entities, of which 193 are members of 
the United Nations) and 21,116 ASNs. If you look with your eyes, you can 
plainly see that a good fraction of those ASNs account for residential 
Internet service.

Given the data, it is not unreasonable to conclude that some kind of 
proxy service that terminates traffic through residential addresses is 
being used. Security researcher Brian Krebs has been tracking this and 
at the very least LG is trying to tackle the problem of people 
installing apps turning their TVs into zombies[2].

If this is all lies, then there sure has been a lot of effort by someone 
to make everyone believe this instead of the more likely explanation: 
it's actually happening.

> There is also an important distinction between seeing a large number of
> requests and establishing that those requests were the cause of a
> particular infrastructure problem. If the concern is resource exhaustion, I
> would expect to see some indication of what resources were actually being
> exhausted: CPU, database queries, bandwidth, connection counts, response
> times, or whatever the relevant bottleneck was. Without that information it
> is difficult to judge either the scale of the original problem or whether
> the measures being taken are proportionate.

One of the main ways resources are being exhausted is that git blame 
queries can take up gigabytes of ram and consume entire CPU cores to 
answer a single query that links to other pages that will also consume 
gigabytes of ram and consume entire CPU cores.

Based on my experiments with the Anubis honeypot subsystem[4] (created 
to figure out how widespread the problem actually is), every single 
request comes from a separate residential IP address. This makes any 
kind of per-IP rate limiting pointless. All rate limiting approaches 
assume there is some kind of pattern between traffic and its source. If 
there is no pattern beyond "the traffic comes from the Internet", all 
attempts to rate limit or throttle based on IP address pointless.

Hell, storing them in ram or trying to iptables them is also pointless 
as the spread of IP addresses is sparse enough that it will create a 
kernel ram exhaustion scenario. When building my own experimental 
databases to try and at least track previously used residential proxies, 
collapsing entries down to the minimal BGP routable units managed to 
save me over 400Mi in maxmind format on disk and 6Gi of ram when 
building those databases. I am still trying to implement this so that 
there can be less user disruption in general, but this is hard.

> There is a useful counterexample here: LWN appears to have dealt with a
> similar scraping problem by addressing the resource bottlenecks rather than
> relying solely on identifying and blocking the scrapers. The fact that it
> was possible to make the service usable again without necessarily stopping
> the scraping itself seems quite relevant.

I'm sure the Sourceware team would appreciate your contributions in 
removing the bottleneck. Why are you not leading with an invitation to 
contribute to the solution instead of complaining about the solution 
that was already put in place?

A really easy win would be to provide an API, UI, and link-compatible 
drop-in replacement for Bugzilla that isn't a CGI script requiring 
mod_perl to run and actually implements caching in a sensible way. That 
alone would account for a vast majority of the MySQL and Apache load of 
the Sourceware servers. Even better if you can edit the cgit / gitweb 
interface to not link to git blame without authentication.

I will provide the Sourceware team a script to fetch the IP addresses 
for GitHub Actions and a few other CI providers so that they can get 
through with minimal interference. If you know of any CI providers that 
are having issues, please link them so they can be allowlisted.

> I would therefore be cautious about presenting the current measures as an
> established response to a demonstrated “AI scraper” problem. There may well
> be a real underlying problem, but I don't think the evidence currently
> available justifies treating the explanation for that problem, or the
> necessity of the current blocking mechanisms, as settled.
> 
> In particular, I think it is worth distinguishing between “we observed
> traffic we believe to be abusive” and “we have demonstrated that this
> traffic is responsible for our problems and that blocking it in this
> particular way is necessary”. Those are very different claims, and the
> latter needs considerably more evidence.

These two arguments read like they were hallucinated by an AI model.

Please provide the original text used when prompting your model and 
details about which AI setup was used to fabricate the utterances in 
your email.

Xe Iaso
(zi yah-so)
[email protected]

.i la budza pu cusku lu
  <<.i ko snura .i ko kanro
    .i ko panpi .i ko gleki
  >>li'u vau iau

[1]: https://xeiaso.net/notes/2026/check-your-smart-tv/full-table/
[2]: 
https://krebsonsecurity.com/2026/07/lg-to-ban-residential-proxies-from-smart-tv-apps/
[3]: https://en.wikipedia.org/wiki/C10k_problem
[4]: https://anubis.techaro.lol/docs/admin/honeypot/overview

PS: I'm sorry if I come across as angry at all in my email. I genuinely 
want to help make the Internet a better place but I work two full time 
jobs and have been recovering from medical leave earlier this year. I'm 
at least no longer in constant pain, so that's a plus.

PS the second: Sorry Kalamatee, I am new to sending emails to mailing 
lists and assumed that pressing "Reply" in my mail client like I have my 
entire career would result in the mailing list being CC'd too.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.