Use case · Content, media, publishing
Distinguishing AI Crawlers From Human Readers
Every request to your site is one of four things: a person, a traditional search crawler, an AI training crawler, or a retrieval agent fetching your page to answer someone's question in real time. Each has a completely different commercial meaning — and right now, most publishers can't tell them apart.
Why this is suddenly a real problem
Quick answer
AI crawler detection identifies, per request, whether traffic to your site is a human, a declared search crawler, an AI training crawler, or an AI retrieval agent — using IP and network signals rather than trusting the User-Agent header, which unsophisticated crawlers spoof and sophisticated ones simply omit.
Publishers have always dealt with bots, but the mix has changed. AI training crawlers harvest your content in bulk with no connection to any specific reader. Retrieval agents fetch a single page in real time to answer someone's question inside a chat interface — a referral source that may or may not turn into a real visit afterwards. Neither behaves like a human, and neither behaves like the search crawlers your analytics were built to handle.
The volume is no longer trivial. On many publishing sites, AI crawler and agent traffic is now a material share of total requests — large enough to distort pageview counts, advertising rates, media kits, and the editorial decisions made from that data.
How IP Raccoon helps
IP Raccoon maintains the aggregated list of known AI crawlers and agents so you don't have to. Nobody wants to own forty different published IP ranges from every AI lab and track which ones changed this month — that's the whole value of a commodity feed.
Real time
Decide per request
Call the API as the request comes in and get back the category — human, search crawler, AI training crawler, AI retrieval agent, or suspected bot — so you can change what the page does before it's served: a paywall shown differently to a training crawler than to a person, for instance.
Offline enrichment
Clean up what you already collected
No real-time requirement: enrich existing event logs, pageviews or subscription-funnel data against downloaded files. The categories let you segment rather than just delete — human traffic, declared crawlers, AI agents, suspected bot — so nothing gets thrown away that you might need later.
Where this actually shows up in the business
Advertising rates & media kits
Inflated pageviews from crawlers overstate real reach, which either overpromises to advertisers or gets caught later and damages trust.
Editorial decisions
Deciding what to commission next based on traffic that includes bot pageviews means optimising for the wrong audience.
Training-data licensing
Knowing which AI labs are crawling you — and how much — is the starting point for any conversation about licensing your content instead of it being scraped for free.
What to look for in an AI crawler detection feed
- Categorisation, not a binary bot/not-bot flag — you need to treat Googlebot, an AI training crawler and a retrieval agent differently, not the same
- Coverage that's actually maintained — new crawlers and agents appear often enough that a stale list is close to useless within months
- Both delivery modes available: real-time API for in-the-moment decisions, files for backend enrichment against logs you already have
- Evidence behind the verdict, not just a label, so your own team can audit why a request was classified a certain way
$ curl 'https://api.ipraccoon.com/ip/203.0.113.42' \
-H "Authorization: $IP_RACCOON_TOKEN"
{
"ip": "203.0.113.42",
"risks": {
"known_bots": "GPTbot"
}
}See what's actually reading your site
Run your own traffic through IP Raccoon and see the real split between people, search crawlers, AI crawlers and agents.
FAQ