Crawler claims

Anyone can say they're Googlebot.

Every request introduces itself in one line of text, and lying costs nothing. RipTide groups crawlers by the company they claim to come from, then checks the claim against the addresses that company publishes. In one week on our own decoy, every visitor calling itself Googlebot failed the check.

RipTide Research9 min read

The short version

A User-Agent is a name tag anyone can print. RipTide treats it as a claim, then checks it against the IP ranges or reverse DNS the claimed operator publishes: verified, mismatch, or unknown.

  1. 1

    Crawlers are a normal part of internet traffic. Grouping them by the company they claim turns dozens of sessions into a handful of labeled groups, so the real activity stands out.

  2. 2

    Each claim is checked offline against the ranges 11 operators publish for their crawlers, and against forward-confirmed reverse DNS for three more that document it instead.

  3. 3

    A failed check is a finding in its own right. On RipTide's own internet-facing decoy, all three sessions claiming to be Googlebot in one week came from outside Google's ranges.

01Act I · The ordinary world

Two hundred sessions. Thirty-eight name tags.

In the last week of September, RipTide's own internet-facing decoy, a small fictional company's website and API, recorded 200 sessions. Nobody was invited. They showed up anyway, because that is what happens to anything with a public address.

Every one of them introduced itself. HTTP requests carry a User-Agent header, one line of text in which the client says what it is: a browser, a script, a search engine. Across those 200 sessions there were 38 different introductions.

Forty-two sessions said they were crawlers. Twenty-five said they came from OpenAI: GPTBot, OAI-SearchBot, ChatGPT-User. Three said they were Censys, three said Palo Alto Networks, three said LeakIX. And three said they were Googlebot, the crawler most websites are glad to see.

A User-Agent is a name tag at a conference. Anyone can print one. The question is who checks the badge.

For a defender, crawlers are background noise you still want to understand. A real search engine reading your decoy is not an incident. A scanner announcing itself is useful context. Something that only says it's a search engine is a different story entirely. The job is to tell those apart, quickly, without anyone reading 200 sessions by hand.

02Act II · The villain

The cheapest disguise on the internet.

Changing a User-Agent takes one line in any HTTP library. Plenty of sites treat search crawlers kindly, so "Googlebot" is a popular costume. It costs an intruder nothing and it can buy a lot: a quieter log line, a looser rate limit, an analyst who glances at the name and moves on.

That leaves a defender two bad options.

Trust the name tag

Every "Googlebot" is Google.

  • Impostors get filed under a household name and ignored.
  • Your reports say "Google" about traffic Google never sent.
  • The disguise works exactly as the intruder hoped.

Ignore the name tag

Every session is a stranger.

  • Twenty-five OpenAI sessions become twenty-five rows to read.
  • Real crawlers drown out the activity that matters.
  • You throw away a useful hint along with the lie.

RipTide takes a third path. It uses the name tag, and it never believes it. Every crawler grouping carries the label self-reported: it records what the client said about itself, never who it is. Then RipTide checks the badge.

03Act II · The struggle

Three ways to check a badge wrong.

Checking a crawler's claim sounds simple: look up the addresses the operator publishes and see if the visitor is inside them. Building it, we found three ways that simple idea goes wrong.

1. The list that would have verified attackers

Google publishes several lists of addresses its crawlers and fetchers use. One of them covers addresses inside Google's cloud platform that customers' own applications fetch from. Include that list, and an attacker running code on Google's cloud and calling itself Googlebot would come out verified. We left it out on purpose. A check that blesses the impostor is worse than no check.

2. The operators who publish no list

Our first week ended with six of eight crawler groups reading unknown, because only some operators publish their crawler addresses in a machine-readable file. Two of them, Censys and Palo Alto Networks, document their scanner ranges on ordinary web pages, so we transcribed those pages, URL and all. Yandex, Baidu and Huawei document something else: the reverse DNS names their crawlers use.

3. The name anyone can set

Reverse DNS has a catch. The name an IP address points back to is set by whoever controls that address, so an attacker can make their address claim to be crawler.yandex.com. A matching name proves nothing on its own. RipTide only verifies when the name ends in the operator's documented domain and the operator's own DNS resolves that name back to the same address. That two-way check is called forward-confirmed reverse DNS.

Even then, DNS is slow, and an attacker who controls an address controls how slowly its lookup answers. So DNS never runs while a request is being answered or a page is being drawn. Lookups happen in the background in small batches, stop at the first failure, and the next page view shows the answer.

04Act III · The turn

Read the claim, group it, check it.

Here's what happens to every session's User-Agent in the RipTide console.

  1. 01

    Set the tools aside

    Generic HTTP libraries and attack tools (curl, python-requests, Go's HTTP client, Nmap, sqlmap, nuclei, zgrab and others, 32 product names in all) never become a crawler group, even if they also mention a vendor. A script is not a crawler.

  2. 02

    Match the claim

    A plain registry of 57 crawler product names across 29 organizations (GPTBot, ClaudeBot, Googlebot, bingbot, PerplexityBot, CCBot, Applebot and the rest) matches the claim. Different bots and versions from one company land in one group, so GPTBot, OAI-SearchBot and ChatGPT-User all read as OpenAI.

  3. 03

    Fall back to the URL

    A bot that isn't in the registry but names its operator's website gets grouped by that site's domain. A vendor's name is only borrowed through an explicit list of its own domains, so a look-alike domain never gets to call itself OpenAI.

  4. 04

    Check the address

    Each member of a crawler group gets a verdict. Verified: the source address is one the claimed operator lists for its crawlers. Mismatch: a public address outside all of them. Unknown: the operator publishes nothing to check against, or the address can't prove anything.

The console shows the verdict as a badge on every group member, counts each verdict per group, and gives an operation the worst verdict among its sessions, so a mismatch never hides inside a bigger cluster. Exports carry it too: a crawler group in a STIX bundle is marked as a self-reported claim, with its verified, mismatch and unknown counts, and it is never turned into a threat actor.

Crawler groups on RipTide's internet-facing decoy, September 20 to 26, 2026
Claimed operatorWhat it called itselfSessionsVerifiedMismatchUnknown
OpenAIGPTBot, OAI-SearchBot, ChatGPT-User252410
GoogleGooglebot3030
CensysCensysInspect3003
Palo Alto NetworksCortex Xpanse scanner3003
LeakIXl9scan3003
ForestEngineForestEngine3003
AgentTrustAgentTrustBot1001
CloudflareSecurity Center scanner1001
Our own decoy's traffic, read from the console the week the feature shipped: 200 sessions, 42 in crawler groups. The four mismatches (three "Googlebots" and one "ChatGPT-User") all came from one proxy provider's address space, which looks like a single client wearing several costumes. Censys and Palo Alto Networks read unknown that week; their documented ranges were added right after.

Look at the Google row. Every session that called itself Googlebot came from somewhere Google doesn't crawl from. Without the check, that's three lines of harmless search traffic. With it, it's one of the most interesting things that happened all week.

Illustrative. Addresses are from documentation ranges; the verdicts mirror the week above.

05Under the hood

Offline by default, careful when it isn't.

The published ranges ship with the sensor

The address lists come from the operators themselves, as machine-readable files: Google, Microsoft (Bing), OpenAI, Anthropic, Apple, Perplexity, DuckDuckGo, Common Crawl and Ahrefs. A refresh tool fetches them all and writes one snapshot into the build, so checking a claim never touches the network. The refresh is all or nothing: if any list fails to fetch or parse, the old snapshot stays, so no operator silently slides from verified to unknown. Each operator's ranges are merged and sorted, so a check is a binary search.

57

crawler product names in the registry, across 29 organizations

RipTide source, October 2026

11 + 3

operators checked against published ranges, plus three checked with forward-confirmed reverse DNS

RipTide source, October 2026

1,105

address prefixes in Google's published crawler lists alone, in the bundled snapshot

RipTide snapshot, September 26, 2026

Reverse DNS, on a short leash

  • Only operators that document reverse DNS (Yandex, Baidu, Huawei) get it. An operator that publishes a list is always checked against the list.
  • At most eight lookups go out per batch, one batch at a time, after a page has already been served.
  • A verdict is kept for 24 hours. A failed lookup reads unknown and is retried after 15 minutes, not on every page view.
  • Reverse DNS names are compared, never stored or shown.

A model for the bots nobody has heard of yet

New crawlers appear all the time. When a User-Agent clearly calls itself a bot but no rule can place it, licensed installs can ask a language model which organization it names. The answer is only accepted if that name actually appears in the User-Agent text, so the model can't invent an operator. Placements made this way carry their own label, llm-inferred, and a different color in the console. Like DNS, it never runs while a request is being answered: at most 20 new bot families are looked up in the background, and answers are cached for 30 days.

Everything a client says is treated as hostile text

Every string in a claim except its category and labels came from the visitor, so the console never builds HTML from it. User-Agents are cut to 512 characters before they are read, so an oversized one is neither scanned in full nor kept whole.

06The moral

Let them introduce themselves. Then check the badge.

Most security tools face a choice between believing what a visitor says and throwing it away. RipTide keeps the introduction, because it's useful, and labels it for what it is: something the visitor said.

Then it does the one thing an impostor can't easily fake, which is come from the right place. A real crawler passes quietly and gets out of your way. A costume gets flagged, and a flagged costume is often the most honest thing an intruder ever tells you: it wanted to be mistaken for someone else.

For a defender, that turns a week of noise into eight labeled groups and four sessions worth a closer look. That's the point of a decoy: fewer things to read, and better reasons to read them.

Scenes marked as illustrative are composites written to show how the technique works, not a record of a specific customer incident. Canary credentials are non-privileged and exist only for detection. RipTide detects and alerts; it never takes destructive action against anyone's infrastructure.

Questions

Fair questions.

Does RipTide block a crawler that fails the check?

No. Verification is evidence in the console and in exports, not a firewall rule. The decoy keeps answering, and the mismatch becomes part of the session's story.

Does verification slow the decoy down or call out to the internet?

Published ranges ship inside the build, so a range check is a local lookup. Reverse DNS and the model fallback run in the background after a console page is served, in small, capped batches. None of it runs while the decoy is answering a request.

Does "verified" mean RipTide knows who sent the request?

It means the address is one the operator lists for its crawlers, nothing more. The group still carries the self-reported label, and RipTide never turns a claim into an attribution. See observed vs. self-reported for why.

Make contact.

Curious how many of your "Googlebots" are real? Tell us where to look.

Book a briefing

Thirty minutes with the people who built it. Bring your hardest question.

  • Watch a live agent set off a detection
  • Map decoys to your crown jewels
  • Plan a first deployment in one sitting

We use your email only to reply. No newsletter, no list, no sharing.

Keyboard shortcuts

T
Change the theme. Shift+T goes back.
?
Show this list
Esc
Close whatever's open