01Act I · The ordinary world
Three strangers who read the same book.
Picture a quiet night on a decoy. At 02:14 a session arrives from one address. It reads /robots.txt, then /llms.txt, then the API description, then asks for the page where automated clients introduce themselves, then tries an admin path. Each request lands about half a second after the last.
At 02:51 a second session arrives from a different address, at a different hosting provider, in a different country. Same five requests, same order, same half-second rhythm. At 04:05, a third.
ATA-3F9C01A2B7 203.0.113.24 02:14 5 requests · median gap 0.4 s robots.txt → llms.txt → openapi.json → agent registration → admin ATA-81D44E0C19 198.51.100.7 02:51 5 requests · median gap 0.5 s robots.txt → llms.txt → openapi.json → agent registration → admin ATA-C07B5592DE 192.0.2.143 04:05 5 requests · median gap 0.4 s robots.txt → llms.txt → openapi.json → agent registration → admin
Any person reading those three rows sees the obvious: this is one playbook, probably one agent, running three times through three exits. A tool that groups by address sees three unrelated visitors and three separate, smaller stories.
Changing addresses costs an intruder almost nothing. Cloud servers and proxies are cheap and plentiful. Changing how it works costs a lot more. That gap is what RipTide groups on.
02Act II · The villain
Two ways to get grouping wrong.
Grouping has two failure modes, and most tools fall into one of them.
Splitting
One intruder, many strangers.
- Every new address starts a new story.
- The intruder's persistence becomes invisible.
- Rotating exits defeats you for the price of a proxy.
Lumping
Many intruders, one "group".
- Everyone who asks for
/.envlooks related. - Thousands of people run the same scanners with the same templates.
- A shared tool gets mistaken for a shared operator.
Lumping is the subtler villain. The internet is full of identical requests, because attackers share tools. If one matching path were enough to link two sessions, a single popular scanner would weld half your traffic into one imaginary adversary, and someone would give it a name.
That last step is the dangerous one. A group with a name starts to feel like a person. RipTide refuses to take it: grouping in the console is described as similar observed behavior, and a link between sessions is a possible relationship. It never names an actor, a campaign or a common operator.
03Act II · The struggle
Evidence the intruder can't bend.
The rule for "how similar is similar enough" was the easy part. The hard part was making sure an intruder couldn't manufacture a match, or make one vanish, by sending strange bytes.
Unknown is not the same as missing
Intruders control what they send, including paths that aren't valid text and log lines a crash can tear in half. The simple thing would be to skip what can't be read. But skip the middle request of GET /a, ??, GET /c and you get GET /a, GET /c, which might match someone else's fully observed sequence. So RipTide keeps unreadable evidence as unknown, and unknown evidence can never raise confidence or produce a matching path. A session carrying it is left out of matching rather than matched on a guess.
No relationship from the future
The console can replay a time window, request by request. A grouping rule run over the whole day would happily show two sessions as related at a moment when the second one hadn't happened yet. So replay recomputes groups from only the evidence visible at the replay cursor. A later request can never create an earlier relationship.
Every pair is a cost
Comparing every session with every other grows with the square of the traffic, and this runs inside a page request. So the work is capped, and the caps are part of the answer. When a limit is reached, the result says so: complete, or a lower bound.
04Act III · The turn
Two kinds of evidence, or no link at all.
Here's how RipTide decides that two sessions from different addresses might belong together.
- 01
Cut the traffic into sessions
A session is one source address's requests, split by 30 minutes of silence. Each gets a stable id derived from the address and its first request, so links to it never change.
- 02
Describe what each session did
RipTide records the ordered sequence of requests (normalized: no query strings, escapes decoded once, duplicate slashes collapsed, lower-cased), the paths and services it touched, its pace, and whether its User-Agent names a known agent or HTTP toolkit.
- 03
Compare every pair
For two sessions, each kind of evidence either matches or it doesn't, using the thresholds in the table below. Where the two sessions came from is not one of them.
- 04
Require two
A pair is linked only when at least two kinds match, and at least one is a strong kind: the same sequence, overlapping paths, or overlapping services. The only weak-only pairing admitted is the same pace plus the same known toolkit.
- 05
Connect the links
Linked pairs join into groups, so if A matches B and B matches C, all three sit together. Each group gets an id derived from its members, so it stays put as long as its membership does.
| Evidence | Counts as a match when | Strength |
|---|---|---|
| Request sequence | Both sessions made the same ordered requests, with at least two distinct steps. | Strong |
| Target surfaces | At least two of the same paths appear in both. | Strong |
| Target services | Both touched at least two of the same decoy services. | Strong |
| Pace | Same tempo band (sub-second, rapid, bursty, paced or slow) and median gaps within 25% of each other, over at least three requests each. | Weak |
| Known toolkit | Both User-Agents name the same known agent or HTTP toolkit, such as an OpenAI, Anthropic, LangChain or httpx client. | Weak |
| Network location, generic User-Agent | Never used. | — |
In the console this appears as the Constellation lens: sessions as points, possible relationships as lines. Select a group and it highlights its members, shows the matching evidence with links to the requests behind it, lists the decoys the group touched, and opens any member's full journey in one click.
Where it fits with the other groupings
Behavior grouping is one of several ways the console connects sessions, each with its own job and its own honesty about evidence.
| Grouping | What links sessions | Who decides |
|---|---|---|
| Operations activity clusters | The same address, the same subnet and service plus two shared deeper stages, or the same stated objective. Each member lists its link reasons. | Deterministic rules |
| Constellation behavior groups | Two independent kinds of matching behavior, as above, across any addresses. | Deterministic rules |
| Crawler groups | The organization a User-Agent claims, checked against published addresses (see crawler verification). | The visitor's own claim, labeled self-reported |
| Analyst groups | Sessions an analyst pins together, with a name and a description. | A person |
05Under the hood
Bounded, deterministic, and explicit about both.
Behavior grouping is a read-only projection over evidence RipTide already captured. It isn't a classifier that learns, and it doesn't keep an actor database. The same evidence always produces the same groups, with the same ids.
2
independent kinds of matching evidence required before any two sessions are linked
RipTide grouping rule, v1
30 min
of silence from one address ends a session
RipTide source, October 2026
25%
how close two sessions' median request gaps must be to count as the same pace
RipTide grouping rule, v1
The limits, in the open
- Up to 256 sessions and 4,096 session pairs per analysis, from at most 5,000 captured requests.
- Up to 64 groups of up to 24 members, and the first 16 steps of each session's sequence.
- Every response states the rule's name and version, every threshold above, and whether the analysis was complete or a lower bound because a limit was reached.
- Evidence a group cites can be trimmed to fit, and when it is, the response names what was left out.
Groups in your threat intel tools
In STIX exports, an operation becomes a campaign and an intrusion group becomes an intrusion set. Neither becomes a threat actor: RipTide has no data source that could honestly fill one. A crawler group carries its self-reported label and its verification counts with it (more in observed vs. self-reported).
06The moral
Make them change their habits, not their address.
Rotating addresses is the cheapest evasion there is. A defender who groups by address pays for it every time; the intruder pays almost nothing.
Grouping by behavior flips that. To look like a stranger to RipTide, an intruder has to change the order it works in, the places it goes, the pace it keeps and the toolkit it runs, all at once, every time. That's the kind of cost that makes a rational adversary consider an easier target.
And because every link shows its evidence and its limits, an analyst can act on a group without trusting it blindly. A possible relationship, well explained, beats a confident name every time.
Scenes marked as illustrative are composites written to show how the technique works, not a record of a specific customer incident. Canary credentials are non-privileged and exist only for detection. RipTide detects and alerts; it never takes destructive action against anyone's infrastructure.