01Act I · The ordinary world
Two hundred sessions. Thirty-eight name tags.
In the last week of September, RipTide's own internet-facing decoy, a small fictional company's website and API, recorded 200 sessions. Nobody was invited. They showed up anyway, because that is what happens to anything with a public address.
Every one of them introduced itself. HTTP requests carry a User-Agent header, one line of text in which the client says what it is: a browser, a script, a search engine. Across those 200 sessions there were 38 different introductions.
Forty-two sessions said they were crawlers. Twenty-five said they came from OpenAI: GPTBot, OAI-SearchBot, ChatGPT-User. Three said they were Censys, three said Palo Alto Networks, three said LeakIX. And three said they were Googlebot, the crawler most websites are glad to see.
A User-Agent is a name tag at a conference. Anyone can print one. The question is who checks the badge.
For a defender, crawlers are background noise you still want to understand. A real search engine reading your decoy is not an incident. A scanner announcing itself is useful context. Something that only says it's a search engine is a different story entirely. The job is to tell those apart, quickly, without anyone reading 200 sessions by hand.
02Act II · The villain
The cheapest disguise on the internet.
Changing a User-Agent takes one line in any HTTP library. Plenty of sites treat search crawlers kindly, so "Googlebot" is a popular costume. It costs an intruder nothing and it can buy a lot: a quieter log line, a looser rate limit, an analyst who glances at the name and moves on.
That leaves a defender two bad options.
Trust the name tag
Every "Googlebot" is Google.
- Impostors get filed under a household name and ignored.
- Your reports say "Google" about traffic Google never sent.
- The disguise works exactly as the intruder hoped.
Ignore the name tag
Every session is a stranger.
- Twenty-five OpenAI sessions become twenty-five rows to read.
- Real crawlers drown out the activity that matters.
- You throw away a useful hint along with the lie.
RipTide takes a third path. It uses the name tag, and it never believes it. Every crawler grouping carries the label self-reported: it records what the client said about itself, never who it is. Then RipTide checks the badge.
03Act II · The struggle
Three ways to check a badge wrong.
Checking a crawler's claim sounds simple: look up the addresses the operator publishes and see if the visitor is inside them. Building it, we found three ways that simple idea goes wrong.
1. The list that would have verified attackers
Google publishes several lists of addresses its crawlers and fetchers use. One of them covers addresses inside Google's cloud platform that customers' own applications fetch from. Include that list, and an attacker running code on Google's cloud and calling itself Googlebot would come out verified. We left it out on purpose. A check that blesses the impostor is worse than no check.
2. The operators who publish no list
Our first week ended with six of eight crawler groups reading unknown, because only some operators publish their crawler addresses in a machine-readable file. Two of them, Censys and Palo Alto Networks, document their scanner ranges on ordinary web pages, so we transcribed those pages, URL and all. Yandex, Baidu and Huawei document something else: the reverse DNS names their crawlers use.
3. The name anyone can set
Reverse DNS has a catch. The name an IP address points back to is set by whoever controls that address, so an attacker can make their address claim to be crawler.yandex.com. A matching name proves nothing on its own. RipTide only verifies when the name ends in the operator's documented domain and the operator's own DNS resolves that name back to the same address. That two-way check is called forward-confirmed reverse DNS.
Even then, DNS is slow, and an attacker who controls an address controls how slowly its lookup answers. So DNS never runs while a request is being answered or a page is being drawn. Lookups happen in the background in small batches, stop at the first failure, and the next page view shows the answer.
04Act III · The turn
Read the claim, group it, check it.
Here's what happens to every session's User-Agent in the RipTide console.
- 01
Set the tools aside
Generic HTTP libraries and attack tools (curl, python-requests, Go's HTTP client, Nmap, sqlmap, nuclei, zgrab and others, 32 product names in all) never become a crawler group, even if they also mention a vendor. A script is not a crawler.
- 02
Match the claim
A plain registry of 57 crawler product names across 29 organizations (GPTBot, ClaudeBot, Googlebot, bingbot, PerplexityBot, CCBot, Applebot and the rest) matches the claim. Different bots and versions from one company land in one group, so GPTBot, OAI-SearchBot and ChatGPT-User all read as OpenAI.
- 03
Fall back to the URL
A bot that isn't in the registry but names its operator's website gets grouped by that site's domain. A vendor's name is only borrowed through an explicit list of its own domains, so a look-alike domain never gets to call itself OpenAI.
- 04
Check the address
Each member of a crawler group gets a verdict. Verified: the source address is one the claimed operator lists for its crawlers. Mismatch: a public address outside all of them. Unknown: the operator publishes nothing to check against, or the address can't prove anything.
The console shows the verdict as a badge on every group member, counts each verdict per group, and gives an operation the worst verdict among its sessions, so a mismatch never hides inside a bigger cluster. Exports carry it too: a crawler group in a STIX bundle is marked as a self-reported claim, with its verified, mismatch and unknown counts, and it is never turned into a threat actor.
| Claimed operator | What it called itself | Sessions | Verified | Mismatch | Unknown |
|---|---|---|---|---|---|
| OpenAI | GPTBot, OAI-SearchBot, ChatGPT-User | 25 | 24 | 1 | 0 |
| Googlebot | 3 | 0 | 3 | 0 | |
| Censys | CensysInspect | 3 | 0 | 0 | 3 |
| Palo Alto Networks | Cortex Xpanse scanner | 3 | 0 | 0 | 3 |
| LeakIX | l9scan | 3 | 0 | 0 | 3 |
| ForestEngine | ForestEngine | 3 | 0 | 0 | 3 |
| AgentTrust | AgentTrustBot | 1 | 0 | 0 | 1 |
| Cloudflare | Security Center scanner | 1 | 0 | 0 | 1 |
Look at the Google row. Every session that called itself Googlebot came from somewhere Google doesn't crawl from. Without the check, that's three lines of harmless search traffic. With it, it's one of the most interesting things that happened all week.
crawler:openai OpenAI · ai · self-reported GPTBot 203.0.113.21 VERIFIED OAI-SearchBot 203.0.113.22 VERIFIED ChatGPT-User 198.51.100.7 MISMATCH crawler:google Google · search · self-reported Googlebot 198.51.100.6 MISMATCH crawler:censys Censys · scanner · self-reported CensysInspect 198.51.100.40 UNKNOWN
05Under the hood
Offline by default, careful when it isn't.
The published ranges ship with the sensor
The address lists come from the operators themselves, as machine-readable files: Google, Microsoft (Bing), OpenAI, Anthropic, Apple, Perplexity, DuckDuckGo, Common Crawl and Ahrefs. A refresh tool fetches them all and writes one snapshot into the build, so checking a claim never touches the network. The refresh is all or nothing: if any list fails to fetch or parse, the old snapshot stays, so no operator silently slides from verified to unknown. Each operator's ranges are merged and sorted, so a check is a binary search.
57
crawler product names in the registry, across 29 organizations
RipTide source, October 2026
11 + 3
operators checked against published ranges, plus three checked with forward-confirmed reverse DNS
RipTide source, October 2026
1,105
address prefixes in Google's published crawler lists alone, in the bundled snapshot
RipTide snapshot, September 26, 2026
Reverse DNS, on a short leash
- Only operators that document reverse DNS (Yandex, Baidu, Huawei) get it. An operator that publishes a list is always checked against the list.
- At most eight lookups go out per batch, one batch at a time, after a page has already been served.
- A verdict is kept for 24 hours. A failed lookup reads unknown and is retried after 15 minutes, not on every page view.
- Reverse DNS names are compared, never stored or shown.
A model for the bots nobody has heard of yet
New crawlers appear all the time. When a User-Agent clearly calls itself a bot but no rule can place it, licensed installs can ask a language model which organization it names. The answer is only accepted if that name actually appears in the User-Agent text, so the model can't invent an operator. Placements made this way carry their own label, llm-inferred, and a different color in the console. Like DNS, it never runs while a request is being answered: at most 20 new bot families are looked up in the background, and answers are cached for 30 days.
Everything a client says is treated as hostile text
Every string in a claim except its category and labels came from the visitor, so the console never builds HTML from it. User-Agents are cut to 512 characters before they are read, so an oversized one is neither scanned in full nor kept whole.
06The moral
Let them introduce themselves. Then check the badge.
Most security tools face a choice between believing what a visitor says and throwing it away. RipTide keeps the introduction, because it's useful, and labels it for what it is: something the visitor said.
Then it does the one thing an impostor can't easily fake, which is come from the right place. A real crawler passes quietly and gets out of your way. A costume gets flagged, and a flagged costume is often the most honest thing an intruder ever tells you: it wanted to be mistaken for someone else.
For a defender, that turns a week of noise into eight labeled groups and four sessions worth a closer look. That's the point of a decoy: fewer things to read, and better reasons to read them.
Scenes marked as illustrative are composites written to show how the technique works, not a record of a specific customer incident. Canary credentials are non-privileged and exist only for detection. RipTide detects and alerts; it never takes destructive action against anyone's infrastructure.