Decoy replies

Ask it anything. It answers in character.

A script runs out of answers. A language model never does, which is exactly why it's dangerous to put one in a decoy. This is how RipTide lets a model improvise every reply without letting it give the decoy away.

RipTide Research11 min read

The short version

Licensed RipTide installs hand each unscripted request to a large language model that writes the decoy's reply in character, in about a third of a second, and throw away any reply that gives the decoy away.

  1. 1

    A canned decoy runs out of answers; a generated one doesn't. Every path an intruder invents gets a page or an API response that fits it, in the voice of the decoy organization.

  2. 2

    The hard part isn't writing replies. It's stopping the model from telling on itself. Our early replies leaked the prompt's own placeholders and its persona, so every reply now passes a set of checks before it's served.

  3. 3

    In our latest test, generated replies arrived in 346 ms at the median, through the RipTide relay or your own Groq or Cerebras key. Anything that fails gets the plain answer a real server would give.

01Act I · The ordinary world

The request nobody scripted.

Picture an attacker's AI agent a few minutes into a visit to a company's API server. It has already tried the obvious paths: the login page, the docs, the health check. Now it starts to improvise.

GET /api/v2/invoices?status=overdue. Then /api/v2/invoices/export?format=csv. Then a path built from a word it saw in the last response. No scanner's wordlist contains these. The agent made them up, the way a person would, from what it has read so far.

This is where most decoys fall apart. A script has an answer for every question its author thought of. An agent asks the ones nobody thought of, and the script's answer to all of them is the same: a blank 404, or the same canned page again. A human might shrug. An agent notices the pattern, and moves on.

On licensed RipTide installs, every request the decoy has no scripted answer for goes to a large language model with two things: who it is playing (the decoy organization and its web server), and what the visitor just asked. The model writes the reply that server would send.

Illustrative and abridged. The path exists nowhere in the decoy's configuration; the reply was written for it.

The reply looks like the API it claims to be, and it does something a script can't: it gives the agent somewhere to go next. Every next step is another request in your logs, and another chance for the visitor to show what it wants and what it is.

02Act II · The villain

Two villains: speed, and a model that loves to talk.

Putting a language model behind a decoy sounds simple. It has two enemies, and we met both.

The first villain was time

We started with a small open-weight model running on the decoy itself, about half a billion parameters, so nothing would leave the box. On a laptop it was quick. Then we measured it on the kind of server a decoy actually lives on: a one-vCPU cloud machine costing $6 a month.

64 s

before the first word of a reply, with the half-billion-parameter model on a one-vCPU, 1 GB server

RipTide measurement, September 23, 2026

1.5

tokens (word pieces) per second after that, on the same server

RipTide measurement, September 23, 2026

346 ms

median time for a complete generated reply today, through the RipTide relay

RipTide measurement, October 2, 2026

The first two numbers are from a calibration run on a DigitalOcean s-1vcpu-1gb server. The third is the median of 14 generated replies on a fresh 1 GB server, measured by the sensor itself.

No real web server takes a minute to answer. We tried to work around it: generating replies ahead of time on a fast laptop and shipping them as a file (we "baked" 374 of them), and a mode that answers a slow request with a plain fallback while the real reply is written in the background for next time. Both helped. Neither fixed the problem. The fix was to move the writing to fast hosted inference, which is what licensed installs use today.

The second villain was the model itself

Language models are trained to be helpful and to explain themselves. A decoy needs the opposite. One reply that sounds like a chatbot, repeats its instructions or stops halfway through a sentence tells the intruder exactly what it's talking to. It only takes one.

A decoy that quotes its own instructions is a decoy that confesses.

03Act II · The struggle

Four times the model told on itself.

Here is the honest list. Each one was caught in our own testing or review, and each one became a permanent check.

  1. 01

    The placeholder (September 21)

    Our first end-to-end run sent six requests to model-backed paths. Four replies came back carrying the prompt's scaffolding instead of real content. One was served with a content type of <mime>, copied literally from our format example. Another was an HTML page whose title was the label we had used to mark the attacker's request as untrusted.

  2. 02

    The persona heading (September 24)

    After the move to hosted inference, 6 of 10 fresh replies were the same 100-byte page: a single heading that restated the description we had given the model of the company and its web server. The same page, on different paths, on two different hosts, is a fingerprint. We rewrote the prompt so nothing copyable sat where the reply should go, and added a check that rejects a reply that is only the persona. The re-test: 0 of 10 echoes, and 10 of 10 replies different from each other.

  3. 03

    The cut-off reply (September 24)

    About a quarter of fresh replies ran out of room at the token limit: HTML that ended mid-tag, JSON that stopped at a key with no value. Real servers don't do that. A cut-off reply is now never served; the visitor gets the plain fallback instead, and we raised the budget from 400 to 512 tokens. Some still run long (4 of 18 in our latest test). They get the plain answer, which is the right way to fail.

  4. 04

    The cached instruction (October 2)

    A review of a stored reply file found 135 of 380 cached replies repeating a sentence of the prompt or the untrusted-request label. None could be served by the live configuration (they were left over from an older setup), but the prompt still contained a sentence a model could copy. Now every instruction sentence in the prompt is on a reject list, a test fails if someone adds a sentence without listing it, and the cache refuses to load any stored reply that fails the check.

Notice the pattern. Every failure was a reply that told the visitor something about the decoy instead of about the server it pretends to be. So the rule we ended up with is simple: the model writes, but it never gets the last word. Code does.

04Act III · The turn

How a request becomes a reply.

Every unscripted request on a licensed install goes through the same five steps.

  1. 01

    Fence the request

    The model sees the method, the path, the query, four headers (Host, User-Agent, Accept, Content-Type), the names of any cookies but never their values, whether an Authorization header was sent but never what it said, and the first 512 bytes of the body. All of it is wrapped and labeled as untrusted data. Anything in the request that tries to close that wrapper early is neutralized before the model sees it.

  2. 02

    Cast the role

    The model is told who it is (the decoy organization and its web server) and asked for exactly three things on three lines: a status, a content type, and a body. Models follow a plain line format more reliably than JSON, and it is easy to check.

  3. 03

    Write the reply

    By default the request goes to the RipTide relay, which answers from Cerebras first and from Groq if Cerebras fails. The provider keys live in the relay, never on your sensor. If you'd rather use your own account, set your own Cerebras or Groq key and the sensor calls it directly. Today both run an open-weight 120-billion-parameter model with its reasoning effort turned down, so the reply fits the budget.

  4. 04

    Check everything

    Headers a real server owns (Server, Date, Set-Cookie, Content-Length, Location and the framing headers) are stripped from whatever the model wrote. Words and phrases that would admit what the decoy is are deleted, and so are private-key blocks. Then the reply is rejected outright if it is empty, cut off, echoes a placeholder, the prompt or the persona, narrates about "the request" instead of answering it, or starts with a raw HTTP status line. A JSON body labeled as HTML gets the right content type.

  5. 05

    Serve it, and remember it

    The decoy's server personality adds its own headers, so the reply sounds like the right server. The reply is cached for an hour, so asking the same question twice gets the same answer; a page that changes on every reload is its own tell. Anything rejected along the way gets the route's plain fallback, a 404 on the decoy website.

What an analyst sees for each exchange. Values are illustrative; the field names are the ones the sensor records.

Every exchange records which provider and model answered, how long it took, whether it came from the cache, and, when the plain fallback answered instead, exactly why. An analyst can always tell a written reply from a fallback. The intruder can't.

05Under the hood

Guard rails, in numbers.

A model in the request path is a dependency, and dependencies fail. These are the limits that keep a slow, busy or broken model from ever hanging or crashing the decoy.

Limits and fallbacks for model-written replies on licensed RipTide installs
GuardSettingWhat happens
Reply budget512 tokensA reply that runs past it is never served. The plain fallback answers.
Provider timeout2.5 s per provider at the relayA slow Cerebras answer fails over to Groq.
Sensor timeout6 s by default for the call to the relay or providerThe sensor stops waiting and serves the plain fallback.
Circuit breakerOpens after 3 timeouts or errors in a rowNo more calls for 5 minutes, then one trial request. Everything in between gets the plain fallback instantly.
Request capsPer minute and per day on the sensor; per license per day at the relayOver a cap, the plain fallback answers. The sensor's own caps never trip the breaker.
Off switch--llm offNo model calls at all. Every route answers with its plain fallback.
Settings as shipped in the licensed profile and the RipTide relay, October 2026.

What a failure looks like from outside

In our latest test we revoked a license key mid-run. For up to three minutes the relay kept serving it, because the relay remembers a key's status for 180 seconds. After that the relay refused, the sensor wrote a single log line, its breaker opened, and every request got the plain 404. No crash, no hang, no restart needed. Restarted with the revoked key, the sensor came back as a trial install and answered from its local reply selector instead.

What leaves the sensor, and what doesn't

  • Only unscripted requests go to a model, and only the fenced summary described above: cookie names but never their values, and never the Authorization header.
  • The relay doesn't log request or response bodies. It logs a provider name, a short failure reason, a duration, and a short prefix of the license key's hash.
  • A license key is only ever sent to the RipTide relay's own host, and the sensor never follows a redirect with it.
  • Installs without a license key send no requests to any model. They choose replies locally.

Why the cache matters more than it looks

The cache is keyed on the service, method, path, query, request body, persona, token budget and model. Headers are left out on purpose, so two different clients asking the same question get the same answer, like they would from a real server. It also means a repeat request costs nothing and arrives in under a millisecond.

06The moral

Stay in character, or stay quiet.

It is easy to judge a model-backed decoy by its best replies: the invoice export that looks real, the admin page with a plausible next link. The replies that matter more are the ones you never see, because they were thrown away.

That is the whole design. The model improvises. Code decides what ships. And when the code isn't sure, the decoy says what a boring real server would say. A plain 404 has never given a decoy away.

For defenders, the result is a decoy with an answer for every question an intruder can invent, which keeps it talking, and keeps it in your logs, for as long as it is willing to stay.

Scenes marked as illustrative are composites written to show how the technique works, not a record of a specific customer incident. Canary credentials are non-privileged and exist only for detection. RipTide detects and alerts; it never takes destructive action against anyone's infrastructure.

Questions

Fair questions.

Does request data go to a third party?

On licensed installs, a fenced summary of each unscripted request (method, path, query, four headers, cookie names and the first 512 bytes of the body) goes to the RipTide relay, which asks Cerebras or Groq to write the reply. The relay doesn't log bodies. With your own Cerebras or Groq key, the summary goes straight to your account. Installs without a license key send nothing to a model.

Can an intruder prompt-inject the decoy?

They can try. Their request reaches the model only as labeled, untrusted data, and attempts to break out of that label are neutralized first. Whatever the model writes is then checked before it is served. A reply that fails a check never reaches the intruder; the plain fallback does.

What if the relay is slow or down?

The relay fails over from Cerebras to Groq. If the sensor still gets no good reply in time, it serves the route's plain fallback, and after three timeouts or errors in a row it stops calling for five minutes. The decoy never hangs waiting for a model.

Make contact.

Send us the strangest path you've seen in your logs. We'll show you what the decoy says back.

Book a briefing

Thirty minutes with the people who built it. Bring your hardest question.

  • Watch a live agent set off a detection
  • Map decoys to your crown jewels
  • Plan a first deployment in one sitting

We use your email only to reply. No newsletter, no list, no sharing.

Keyboard shortcuts

T
Change the theme. Shift+T goes back.
?
Show this list
Esc
Close whatever's open