01Act I · The ordinary world
At 3:07 a.m., something asks for /actuator/env.
Picture a small web server for a company called Acme. It has a homepage, a login page and an API. It does not run Spring Boot. It never has.
At 3:07 in the morning, a request arrives anyway: GET /actuator/env. A second later, /wp-login.php. Then /.env.production, /api/tags, /.git/config. Nobody typed these. A scanner did, or an attacker's AI agent working down a list of things that are often left open on servers like this one.
Every server on the internet hears these questions all day. A normal server gives the honest answer: 404 Not Found. For a normal server, that's exactly right.
For a decoy, it's a missed opportunity. A 404 tells the visitor to move along, and it does. You learn one line of log and nothing about what it wanted next.
A good librarian doesn't write a new book for every question. They walk to the right shelf.
RipTide's decoys answer instead. Ask for the Spring actuator and you get a Spring actuator. Ask for the WordPress login and you get a WordPress login, branded with the decoy organization's name. The visitor keeps going, and every step it takes is another line of evidence: what it was looking for, what it did with what it found, and whether a person or a machine is behind it.
$ curl -i http://203.0.113.10/actuator/env
HTTP/1.1 200 OK
Server: Apache/2.4.58 (Ubuntu)
Content-Type: application/vnd.spring-boot.actuator.v3+json
{"activeProfiles":["prod"],"propertySources":[ …
"spring.application.name":{"value":"acme-api"}, …
"spring.datasource.password":{"value":"******"}, … ]}
That's the job. The hard part is doing it well on the kind of machine people actually give a free trial: a spare VM, a $4 cloud server, a lab box with no GPU.
02Act II · The villain
Writing every reply was too heavy for the smallest boxes.
There are two classic ways a decoy answers a request it wasn't scripted for, and both have a villain hiding in them.
Canned responses
Cheap, and easy to spot.
- The same 404, or the same page, for every unknown path.
- Scanners and agents notice a server that answers everything the same way.
- Once a decoy is recognized, the intruder simply avoids it.
Generate every reply
Believable, and expensive.
- A language model writes a fresh, plausible response to any request.
- In the cloud, every reply costs tokens, needs a network path and an account.
- Run locally, the model is the biggest thing in the binary.
RipTide's paid tier generates replies with a large model in the cloud (more on that in its own story). For the free trial we wanted something that runs entirely on the sensor, costs nothing per request, and needs no network to answer.
Our first answer was to ship a small generative model inside the binary: Qwen2.5-0.5B, an open-weight model of about half a billion parameters. It worked. It also weighed 491 MB, which made the whole binary 505 MB. And the launcher that unpacks RipTide reads its entire payload into memory before anything starts.
So the question changed. Not "how do we write a good reply on a small box?" but "do we need to write anything at all?"
Look at what attackers ask for and the answer is mostly no. The questions are predictable. WordPress, phpMyAdmin, .git folders, .env files, Spring actuators, Jenkins, Kubernetes, cloud metadata, Ollama, MCP, Log4Shell probes. Most of the time a request doesn't need a new reply. It needs the right one.
03Act II · The struggle
We bet on a clever model. It lost.
Choosing among options is a different job from writing text, and there are models built for exactly that. Laya, an open-source decision model from Convai Innovations, takes a question and a list of options and returns its choice with a confidence score. On paper it was a perfect fit: give it the request, give it the list of replies, serve whatever it picks.
We didn't take the paper's word for it. We wrote 79 realistic attacker requests by hand (71 matched to the reply they should get, 8 that match nothing and should fall through), and ran every candidate on the machine a trial customer would actually use: a fresh one-vCPU, 1 GB cloud server with no swap, with the RipTide decoy running beside it the whole time.
It didn't go well.
- 01
Memory
Every Laya model loaded the normal way was killed by the kernel for running out of memory, at about 700 MB. The only variant that loaded at all, with its weights memory-mapped from disk, left 82 MiB free for the entire server.
- 02
Speed
The configurations that fit took 2.8 seconds (English) and 5.6 seconds (multilingual) per request at the median. A real server answers in milliseconds. A decoy that pauses for five seconds is a decoy that gets noticed.
- 03
Accuracy
The best Laya run picked the right reply for 62% of requests. Worse, its wrong answers came back confident: on the laptop, wrong picks had a median confidence of 0.94. You can't set a "only answer when sure" threshold on a model that's always sure.
Next to it, almost as an afterthought, we had a control: a small sentence-embedding model, the kind search engines use to tell that "reset my password" and "forgot login" mean nearly the same thing. It doesn't reason about the question. It just measures how close the request is to each reply's description.
| Candidate | Model size | Time per request (median) | Right reply | Free memory left |
|---|---|---|---|---|
| MiniLM-L6 sentence embeddings int8, the one we shipped | 23 MB | 6.9 ms | 76.1% | 631 MiB |
| Laya multilingual int8, weights memory-mapped | 336 MB | 5,606 ms | 62.0% | 82 MiB |
| Laya English int8, weights memory-mapped | 444 MB | 2,752 ms | 43.7% | 209 MiB |
| Laya, loaded normally English or multilingual, int8 | 336–445 MB | Killed by the kernel while loading | — | — |
A model one-fifteenth the size won on every axis we measured: hundreds of times faster, more accurate, and leaving nearly all the memory to the decoy. The lesson stung a little, and we kept it anyway: the clever tool and the right tool are not always the same tool.
04Act III · The turn
How the reply selector works.
What shipped is called the reply selector. It's the default for every RipTide install without a license key, and it works in five steps.
- 01
A library of replies
We wrote 69 replies across 12 families: CMS logins, database tools, secret files, dependency manifests, Java and CI consoles, APIs, cloud and container APIs, AI endpoints, exploit payloads, appliances, admin pages and debug pages. Each one is a complete HTTP answer (status, headers, body) that looks like the decoy organization's real software, plus a short description of what it answers.
- 02
Learn the descriptions once
When RipTide starts, the model turns each description into a list of 384 numbers, a point in "meaning space". That happens once, in under two seconds.
- 03
Read the request
Each unscripted request becomes a short text: the method and path, the query, the headers that say something about intent, and the first 200 characters of any body. Headers every request carries are dropped (more on why below).
- 04
Find the nearest reply
The request goes into the same meaning space, and the selector measures how close it lands to each of the 69 descriptions. That's one small multiplication per reply.
- 05
Answer, or step aside
If the best match scores at least 0.42, the decoy serves that reply. If not, it answers with the same plain 404 the real server would. Either way, the exchange is recorded with which reply answered, its score and how long it took.
Why 0.42? Because of those eight requests that match nothing. The highest any of them ever scored was 0.401. Setting the bar just above it means the selector never answers a request the library doesn't cover. On the labeled set it answered 64 of 79 requests, got 58 of them right, and answered none of the eight it shouldn't have.
08:07:12 GET /actuator/env 200 provider embed (reply selector) entry actuator_env score 0.711 latency 8 ms 08:07:13 GET /about-us 404 fallback below_threshold
05Under the hood
The details that made the difference.
Two ordinary headers cost 15 points of accuracy
The surprise of the whole project was the smallest one. Our hand-written requests didn't include a Host header. Real requests always do, and most send Accept: */*. Adding just those two, the way curl sends them, dropped the selector from 87.3% to 71.8% correct.
The reason is that the model reads the whole request as one piece of text. Lines that every request shares pull every request toward the same place, and the path, which carries the meaning, gets drowned out. Long browser User-Agent strings did the same thing. So the selector now drops headers that appear whatever a request is probing for (Host, Connection, Accept-Encoding, forwarding headers, Sec-* and the like), and drops the User-Agent unless it carries ${, so a Log4Shell probe hidden there still gets read.
Descriptions beat examples
We tried teaching the selector with example requests for each reply, the way you might train a classifier. Accuracy fell to 55–58%. Raw request lines look more like each other than like their meaning. A short, plain description of what each reply answers works better.
The replies are built not to give the decoy away
- Every reply matches the host it claims to be: Apache 2.4.58 on Ubuntu, with matching PHP and MySQL versions, never another server's banner. Tests enforce it.
- Statuses vary on purpose (200, 401, 403), because a server that says yes to every probe is its own fingerprint.
- No reply echoes anything from the request back, so a probe can't inject into a reply.
- Every reply carries the decoy organization's names, so
--var COMPANY=…rebrands all 69 at once. - Credentials in replies are either masked, placeholders, or planted canaries the console already watches for, never real secrets.
What it costs to run
45.7 MB
default binary, down from 505 MB with the bundled generative model
RipTide build, October 2026
12 ms
median time to choose a reply on a one-vCPU, 1 GB server (40 ms at p95)
RipTide measurement, October 2026
512 MB
the smallest cloud server we ran the trial on, start to finish, with no out-of-memory errors
RipTide measurement, October 2026
Measured on fresh DigitalOcean servers (s-1vcpu-1gb and s-1vcpu-512mb-10gb, Ubuntu 26.04) with the full RipTide sensor and console running. The selector is ready about 1.8 seconds after start.
Where it sits next to the other options
| Install | How unscripted requests are answered | Network needed for replies |
|---|---|---|
| Trial no license key | The reply selector picks one of 69 pre-written replies, or answers 404. | None |
| Licensed | A large language model writes a fresh reply through the RipTide relay, or through your own Groq or Cerebras key. | Outbound to the relay or your provider |
| Your own local model | Supply your own GGUF model file and run a profile set to the local provider; RipTide generates replies on the sensor. | None |
06The moral
Small, boring and right.
It would have been a better headline to say we put a reasoning model inside every decoy. It would also have been a worse decoy: slower, heavier and less accurate on the questions attackers actually ask.
A decoy reply doesn't need to be eloquent. It needs to be believable, fast and on topic, so the intruder takes one more step. Every extra step is another request in your logs, another chance for the intruder to show what it wants and what it is. That's the whole game: keep them talking long enough to be seen.
For defenders, it means a free trial that answers like the real thing on the smallest server you have, without a single reply leaving the box. For us, it was a reminder worth writing down: measure on the machine your customer will use, and let the numbers pick the model.
Scenes marked as illustrative are composites written to show how the technique works, not a record of a specific customer incident. Canary credentials are non-privileged and exist only for detection. RipTide detects and alerts; it never takes destructive action against anyone's infrastructure.