Download the latest Vulnerability & Exploitation Report

Download now

The PaperPhone Cluster: One Operator, 80K IPs, 43 countries

On the 31st of August, we released CrowdSec 1.8.0 that included the bot detection feature. Despite the fact that the feature was (is) still tagged as alpha, some mad lads promptly slapped it on their production infrastructure. Thanks for the trust!

This allowed us to progressively collect some bot signal data that we promptly started to investigate. Until the addition of bot detection features within CrowdSec, the Security Engine and Web Application Firewall have been nearly exclusively focused on “what is the IP doing?”, so being able to look into “what is the client behind this IP running?” is for us a whole new perspective, and we’re digging into the data like kids into a bag of unmonitored candies. 

So today, I’m super happy to present to you the first of an hopefully long series of investigations where we discover together the fantastic world of large-scale scrappers that might quickly turn into “What the actual F is m247 up to?”

Yes, we identified a cluster of bots that we dubbed “paperphone”; let me walk you through! But first, the “big numbers” headline:

What CrowdSec bot detection revealed

  • 75,000 IPs
  • Spread across 230 IP blocks
  • Located across 43 countries
  • Spotted over the course of two weeks

The timeline

The timeline itself is to be taken with a grain of salt, as the beginning of the identification of the cluster started when the feature was released: the infrastructure is likely to have been running for a while already.

We can, however, see that the IPs are progressively rallying the cluster as our users are deploying the feature and start to contribute to the detection. Unsurprisingly, we can also see the operator starting to rotate and cycle its IP addresses as they get banned left and right for failing the challenges too often:

But I’m already digressing; let’s look at what this cluster is made of, and how it’s distributed!

Geographic diversity, or just a facade?

Very often “geoblocking” comes into people’s conversations when dealing with unwanted traffic. But as soon as you’re dealing with even remotely sophisticated (emphasis on the remotely) actors, it becomes very ineffective, as “it’s coming from everywhere!”

In this specific case, the traffic from the cluster is originating from 43 distinct countries (if you trust geolocation, but we’ll come back to this later). However, the time of the requests doesn’t align with the time zones of the countries they claim to originate from

CountryPeak local timeNight share
🇯🇵 Japan05:0018.7%
🇮🇳 India01:0018.5%
🇲🇽 Mexico03:0017.4%
🇱🇹 Lithuania02:0016.6%

It goes further than having night-bird users. While the traffic is spread over days, synchronization regardless of the origin countries still shows up. For example, while Japan and the US are 13 hours apart, request spikes from both countries are highly correlated. And this pattern repeats over and over: Australia and Canada (14 hours apart) infrastructure wakes up at the same time.

So, while this seems odd, if you look at it from a UTC time zone point of view, it makes a lot more sense:

Nationality is a field somebody types

The scraper network spreads over 230 distinct IP blocks (mostly /24’s) and 80 networks. But again, this apparent diversity is a mere facade. When looking more closely at the IP blocks, a lot of things start to stand out.

Nearly every block involved shows some inconsistencies.

For example, 103.216.1.0/24 was issued to APNIC (The Regional Internet Registry for Asia Pacific). But it’s registered to RIPE (The RIR for Europe, the Middle East, and Central Asia), meaning an inter-regional transfer happened (it was last transferred in 2025). The registrant itself is Lithuanian, while the country claimed by the WHOIS is “United States”. More concretely, it means that the range changed ownership during its lifetime (lastly in 2025) – which is fine – but more importantly, its registration region, the country of the organization managing it, and the claim geolocalisation all differ significantly.

If you start to look closely within the ranges allocated, you can quickly see that the geolocation is to be taken with a grain of salt, as a lot of adjacent ranges are artificially located in various countries:

  • 62.105.200.0/22 : 🇧🇪 Brussels
  • 62.105.204.0/23 : 🇹🇭 Bangkok 
  • 62.105.208.0/23 : 🇯🇵 Tokyo
  • 62.105.210.0/23 : 🇫🇷 Paris

And this pattern is quite systematic when you look at the 230 address blocks involved. The good news is that it all boils down to a few entities that are controlling a significant portion of the involved IP ranges:

Note: The graph above shows ownership of IP ranges: organisation -> country -> AS -> IP block.

What all this shows is that the entities that are in control of the IP blocks deliberately obscure their geographic footprint, making it harder to detect and block. We’re not facing someone renting machines in 43 countries, but an organization artificially claiming to have a diverse geographic presence.

Surprisingly, the name of M247 doesn’t appear in the graph above. Worry not, while none of the blocks belong to M247, it is still the transit provider for more than 20% of PaperPhone’s cluster.

iPhones, Android, and bot signals

These whopping 75 thousand IPs have been showing a very consistent behavior, meticulously cycling between thirteen different device identities, including 5 different Android devices and 7 variants of iOS.

However, within the signals, there were some funny giveaways of obvious lies:

An interesting inconsistency was that the viewport size – the effective visible rendering area – was 375×812 for 100% of the bots. This matches the iPhone 10/11 (to my knowledge) viewport size, inconsistent with the claimed devices.

All of them are also using Google SwiftShader as their WebGL renderer (UNMASKED_RENDERER_WEBGL), which is Google’s open-source CPU-based software renderer. This information means that the Chrome-based browser is running on hardware that doesn’t provide any hardware acceleration.

The claimed Android devices (Pixel 9, Samsung Galaxy S25 Ultra, etc.) are all fairly recent Android devices, so the likelihood of them not having hardware acceleration is approximately zero.

At the risk of stating the obvious, when it comes to iOS 14/15 triggering a SwiftShader signal, given that it’s running Apple’s WebKit, the likelihood of it happening is again zero.

Maxing out available ranges

In this specific case, the IP ranges were mostly already flagged as being data centers and “known” in our CTI, so we’re not facing the more sneaky case of residential proxy abuse, as the density of IP addresses also shows:

Note: The graphic above shows the percentage of IPs in each block that were involved in the scraping actions.

We can see that more of the blocks have been used to their full capacity during the scraping operation, and that all the “diversity” of the network (AS, ranges, countries) is a mere facade.

WRITTEN BY

You may also like

CrowdSec Statement: Source Code Exposure in May 2026
Announcement

CrowdSec Statement: Source Code Exposure in May 2026

CrowdSec update on a source code exposure that occurred in May 2026, including the scope, impact, investigation, and security measures taken.

How to Block Bots, Headless Browsers, and Scrapers on Nginx with CrowdSec
AI

How to Block Bots, Headless Browsers, and Scrapers on Nginx with CrowdSec

Set up open-source Nginx bot protection with CrowdSec to block scrapers, headless browsers, AI crawlers, and other unwanted automated traffic.

What’s New in CrowdSec 1.8: WAF Bot Detection, Kubernetes Datasource, and Performance Improvements
AI

What’s New in CrowdSec 1.8: WAF Bot Detection, Kubernetes Datasource, and Performance Improvements

Discover what’s new in CrowdSec 1.8, including WAF bot detection, a dedicated Kubernetes datasource, faster LAPI synchronization, and improved Console alerts.