Decoy replies

The reply that's chosen, not written.

A decoy that answers "404" is a door closing. RipTide's free trial answers like the real thing instead, and it does it with a 23 MB model that chooses a reply rather than writing one. Here's how we got there, including the clever model that didn't make it.

RipTide Research11 min read

The short version

Instead of writing a reply with a big model, RipTide's trial chooses one from 69 pre-written replies with a 23 MB local model, in about 12 milliseconds on a one-vCPU server.

  1. 1

    A decoy should answer the questions attackers actually ask. Most of those questions are predictable: WordPress logins, .env files, Spring actuators, model servers.

  2. 2

    We measured a purpose-built decision model, Laya, on a 1 GB server. It ran out of memory or took seconds per request. A 23 MB sentence-embedding model did the job better in milliseconds.

  3. 3

    The result: the default binary went from 505 MB to 45.7 MB, starts on a $4 512 MB server, and answers without a GPU, an API bill or any network call for its replies.

01Act I · The ordinary world

At 3:07 a.m., something asks for /actuator/env.

Picture a small web server for a company called Acme. It has a homepage, a login page and an API. It does not run Spring Boot. It never has.

At 3:07 in the morning, a request arrives anyway: GET /actuator/env. A second later, /wp-login.php. Then /.env.production, /api/tags, /.git/config. Nobody typed these. A scanner did, or an attacker's AI agent working down a list of things that are often left open on servers like this one.

Every server on the internet hears these questions all day. A normal server gives the honest answer: 404 Not Found. For a normal server, that's exactly right.

For a decoy, it's a missed opportunity. A 404 tells the visitor to move along, and it does. You learn one line of log and nothing about what it wanted next.

A good librarian doesn't write a new book for every question. They walk to the right shelf.

RipTide's decoys answer instead. Ask for the Spring actuator and you get a Spring actuator. Ask for the WordPress login and you get a WordPress login, branded with the decoy organization's name. The visitor keeps going, and every step it takes is another line of evidence: what it was looking for, what it did with what it found, and whether a person or a machine is behind it.

Illustrative, abridged. The decoy runs Apache on Ubuntu like the rest of the site, and the reply carries the organization's own names.

That's the job. The hard part is doing it well on the kind of machine people actually give a free trial: a spare VM, a $4 cloud server, a lab box with no GPU.

02Act II · The villain

Writing every reply was too heavy for the smallest boxes.

There are two classic ways a decoy answers a request it wasn't scripted for, and both have a villain hiding in them.

Canned responses

Cheap, and easy to spot.

  • The same 404, or the same page, for every unknown path.
  • Scanners and agents notice a server that answers everything the same way.
  • Once a decoy is recognized, the intruder simply avoids it.

Generate every reply

Believable, and expensive.

  • A language model writes a fresh, plausible response to any request.
  • In the cloud, every reply costs tokens, needs a network path and an account.
  • Run locally, the model is the biggest thing in the binary.

RipTide's paid tier generates replies with a large model in the cloud (more on that in its own story). For the free trial we wanted something that runs entirely on the sensor, costs nothing per request, and needs no network to answer.

Our first answer was to ship a small generative model inside the binary: Qwen2.5-0.5B, an open-weight model of about half a billion parameters. It worked. It also weighed 491 MB, which made the whole binary 505 MB. And the launcher that unpacks RipTide reads its entire payload into memory before anything starts.

So the question changed. Not "how do we write a good reply on a small box?" but "do we need to write anything at all?"

Look at what attackers ask for and the answer is mostly no. The questions are predictable. WordPress, phpMyAdmin, .git folders, .env files, Spring actuators, Jenkins, Kubernetes, cloud metadata, Ollama, MCP, Log4Shell probes. Most of the time a request doesn't need a new reply. It needs the right one.

03Act II · The struggle

We bet on a clever model. It lost.

Choosing among options is a different job from writing text, and there are models built for exactly that. Laya, an open-source decision model from Convai Innovations, takes a question and a list of options and returns its choice with a confidence score. On paper it was a perfect fit: give it the request, give it the list of replies, serve whatever it picks.

We didn't take the paper's word for it. We wrote 79 realistic attacker requests by hand (71 matched to the reply they should get, 8 that match nothing and should fall through), and ran every candidate on the machine a trial customer would actually use: a fresh one-vCPU, 1 GB cloud server with no swap, with the RipTide decoy running beside it the whole time.

It didn't go well.

  1. 01

    Memory

    Every Laya model loaded the normal way was killed by the kernel for running out of memory, at about 700 MB. The only variant that loaded at all, with its weights memory-mapped from disk, left 82 MiB free for the entire server.

  2. 02

    Speed

    The configurations that fit took 2.8 seconds (English) and 5.6 seconds (multilingual) per request at the median. A real server answers in milliseconds. A decoy that pauses for five seconds is a decoy that gets noticed.

  3. 03

    Accuracy

    The best Laya run picked the right reply for 62% of requests. Worse, its wrong answers came back confident: on the laptop, wrong picks had a median confidence of 0.94. You can't set a "only answer when sure" threshold on a model that's always sure.

Next to it, almost as an afterthought, we had a control: a small sentence-embedding model, the kind search engines use to tell that "reset my password" and "forgot login" mean nearly the same thing. It doesn't reason about the question. It just measures how close the request is to each reply's description.

Reply selection candidates measured on a one-vCPU, 1 GB cloud server with the decoy running
CandidateModel sizeTime per request (median)Right replyFree memory left
MiniLM-L6 sentence embeddings int8, the one we shipped23 MB6.9 ms76.1%631 MiB
Laya multilingual int8, weights memory-mapped336 MB5,606 ms62.0%82 MiB
Laya English int8, weights memory-mapped444 MB2,752 ms43.7%209 MiB
Laya, loaded normally English or multilingual, int8336–445 MBKilled by the kernel while loading——
RipTide measurements, October 1–2, 2026, on a DigitalOcean s-1vcpu-1gb server (one vCPU, 956 MiB, no swap) against 71 hand-written, labeled attacker requests. These are synthetic requests, not production traffic. Laya is built for richer decisions than this one; on a box this small, it was the wrong tool.

A model one-fifteenth the size won on every axis we measured: hundreds of times faster, more accurate, and leaving nearly all the memory to the decoy. The lesson stung a little, and we kept it anyway: the clever tool and the right tool are not always the same tool.

04Act III · The turn

How the reply selector works.

What shipped is called the reply selector. It's the default for every RipTide install without a license key, and it works in five steps.

  1. 01

    A library of replies

    We wrote 69 replies across 12 families: CMS logins, database tools, secret files, dependency manifests, Java and CI consoles, APIs, cloud and container APIs, AI endpoints, exploit payloads, appliances, admin pages and debug pages. Each one is a complete HTTP answer (status, headers, body) that looks like the decoy organization's real software, plus a short description of what it answers.

  2. 02

    Learn the descriptions once

    When RipTide starts, the model turns each description into a list of 384 numbers, a point in "meaning space". That happens once, in under two seconds.

  3. 03

    Read the request

    Each unscripted request becomes a short text: the method and path, the query, the headers that say something about intent, and the first 200 characters of any body. Headers every request carries are dropped (more on why below).

  4. 04

    Find the nearest reply

    The request goes into the same meaning space, and the selector measures how close it lands to each of the 69 descriptions. That's one small multiplication per reply.

  5. 05

    Answer, or step aside

    If the best match scores at least 0.42, the decoy serves that reply. If not, it answers with the same plain 404 the real server would. Either way, the exchange is recorded with which reply answered, its score and how long it took.

Why 0.42? Because of those eight requests that match nothing. The highest any of them ever scored was 0.401. Setting the bar just above it means the selector never answers a request the library doesn't cover. On the labeled set it answered 64 of 79 requests, got 58 of them right, and answered none of the eight it shouldn't have.

What an analyst sees for two exchanges. Scores and latency are from a fresh 1 GB test server; the times are illustrative. The selector's choice is recorded next to the raw request.

05Under the hood

The details that made the difference.

Two ordinary headers cost 15 points of accuracy

The surprise of the whole project was the smallest one. Our hand-written requests didn't include a Host header. Real requests always do, and most send Accept: */*. Adding just those two, the way curl sends them, dropped the selector from 87.3% to 71.8% correct.

The reason is that the model reads the whole request as one piece of text. Lines that every request shares pull every request toward the same place, and the path, which carries the meaning, gets drowned out. Long browser User-Agent strings did the same thing. So the selector now drops headers that appear whatever a request is probing for (Host, Connection, Accept-Encoding, forwarding headers, Sec-* and the like), and drops the User-Agent unless it carries ${, so a Log4Shell probe hidden there still gets read.

Descriptions beat examples

We tried teaching the selector with example requests for each reply, the way you might train a classifier. Accuracy fell to 55–58%. Raw request lines look more like each other than like their meaning. A short, plain description of what each reply answers works better.

The replies are built not to give the decoy away

  • Every reply matches the host it claims to be: Apache 2.4.58 on Ubuntu, with matching PHP and MySQL versions, never another server's banner. Tests enforce it.
  • Statuses vary on purpose (200, 401, 403), because a server that says yes to every probe is its own fingerprint.
  • No reply echoes anything from the request back, so a probe can't inject into a reply.
  • Every reply carries the decoy organization's names, so --var COMPANY=… rebrands all 69 at once.
  • Credentials in replies are either masked, placeholders, or planted canaries the console already watches for, never real secrets.

What it costs to run

45.7 MB

default binary, down from 505 MB with the bundled generative model

RipTide build, October 2026

12 ms

median time to choose a reply on a one-vCPU, 1 GB server (40 ms at p95)

RipTide measurement, October 2026

512 MB

the smallest cloud server we ran the trial on, start to finish, with no out-of-memory errors

RipTide measurement, October 2026

Measured on fresh DigitalOcean servers (s-1vcpu-1gb and s-1vcpu-512mb-10gb, Ubuntu 26.04) with the full RipTide sensor and console running. The selector is ready about 1.8 seconds after start.

Where it sits next to the other options

How RipTide answers unscripted requests, by install type
InstallHow unscripted requests are answeredNetwork needed for replies
Trial no license keyThe reply selector picks one of 69 pre-written replies, or answers 404.None
LicensedA large language model writes a fresh reply through the RipTide relay, or through your own Groq or Cerebras key.Outbound to the relay or your provider
Your own local modelSupply your own GGUF model file and run a profile set to the local provider; RipTide generates replies on the sensor.None

06The moral

Small, boring and right.

It would have been a better headline to say we put a reasoning model inside every decoy. It would also have been a worse decoy: slower, heavier and less accurate on the questions attackers actually ask.

A decoy reply doesn't need to be eloquent. It needs to be believable, fast and on topic, so the intruder takes one more step. Every extra step is another request in your logs, another chance for the intruder to show what it wants and what it is. That's the whole game: keep them talking long enough to be seen.

For defenders, it means a free trial that answers like the real thing on the smallest server you have, without a single reply leaving the box. For us, it was a reminder worth writing down: measure on the machine your customer will use, and let the numbers pick the model.

Scenes marked as illustrative are composites written to show how the technique works, not a record of a specific customer incident. Canary credentials are non-privileged and exist only for detection. RipTide detects and alerts; it never takes destructive action against anyone's infrastructure.

Questions

Fair questions.

Isn't a pre-written reply easy to fingerprint?

A reply only goes out when the request fits it; everything else gets the same 404 the real server would send. Each reply is a complete page or document consistent with the host's claimed software, carries your decoy organization's names, and statuses vary. Licensed installs go further and generate a fresh reply for each request.

Does the selector send anything off the box?

No. The model and the reply library ship inside the binary, and choosing happens on the sensor. Trial installs need no network for replies.

What happens to requests the library doesn't cover?

They fall below the 0.42 threshold and get the route's normal fallback, a plain 404. They're still captured and logged like every other request, marked as below threshold, so nothing goes unseen.

Make contact.

Tell us the smallest box you'd want a decoy on. We'll show you it answering.

Book a briefing

Thirty minutes with the people who built it. Bring your hardest question.

  • Watch a live agent set off a detection
  • Map decoys to your crown jewels
  • Plan a first deployment in one sitting

We use your email only to reply. No newsletter, no list, no sharing.

Keyboard shortcuts

T
Change the theme. Shift+T goes back.
?
Show this list
Esc
Close whatever's open