01Act I · The ordinary world
The request nobody scripted.
Picture an attacker's AI agent a few minutes into a visit to a company's API server. It has already tried the obvious paths: the login page, the docs, the health check. Now it starts to improvise.
GET /api/v2/invoices?status=overdue. Then /api/v2/invoices/export?format=csv. Then a path built from a word it saw in the last response. No scanner's wordlist contains these. The agent made them up, the way a person would, from what it has read so far.
This is where most decoys fall apart. A script has an answer for every question its author thought of. An agent asks the ones nobody thought of, and the script's answer to all of them is the same: a blank 404, or the same canned page again. A human might shrug. An agent notices the pattern, and moves on.
On licensed RipTide installs, every request the decoy has no scripted answer for goes to a large language model with two things: who it is playing (the decoy organization and its web server), and what the visitor just asked. The model writes the reply that server would send.
$ curl -i http://203.0.113.20/api/v2/invoices?status=overdue
HTTP/1.1 200 OK
Server: Apache/2.4.58 (Ubuntu)
Content-Type: application/json
{"data":[{"id":"INV-10482","customer":"Northwind Freight",
"amount_due":"18,240.00","due":"2026-09-12", … }],
"next":"/api/v2/invoices?status=overdue&page=2"}
The reply looks like the API it claims to be, and it does something a script can't: it gives the agent somewhere to go next. Every next step is another request in your logs, and another chance for the visitor to show what it wants and what it is.
02Act II · The villain
Two villains: speed, and a model that loves to talk.
Putting a language model behind a decoy sounds simple. It has two enemies, and we met both.
The first villain was time
We started with a small open-weight model running on the decoy itself, about half a billion parameters, so nothing would leave the box. On a laptop it was quick. Then we measured it on the kind of server a decoy actually lives on: a one-vCPU cloud machine costing $6 a month.
64 s
before the first word of a reply, with the half-billion-parameter model on a one-vCPU, 1 GB server
RipTide measurement, September 23, 2026
1.5
tokens (word pieces) per second after that, on the same server
RipTide measurement, September 23, 2026
346 ms
median time for a complete generated reply today, through the RipTide relay
RipTide measurement, October 2, 2026
The first two numbers are from a calibration run on a DigitalOcean s-1vcpu-1gb server. The third is the median of 14 generated replies on a fresh 1 GB server, measured by the sensor itself.
No real web server takes a minute to answer. We tried to work around it: generating replies ahead of time on a fast laptop and shipping them as a file (we "baked" 374 of them), and a mode that answers a slow request with a plain fallback while the real reply is written in the background for next time. Both helped. Neither fixed the problem. The fix was to move the writing to fast hosted inference, which is what licensed installs use today.
The second villain was the model itself
Language models are trained to be helpful and to explain themselves. A decoy needs the opposite. One reply that sounds like a chatbot, repeats its instructions or stops halfway through a sentence tells the intruder exactly what it's talking to. It only takes one.
A decoy that quotes its own instructions is a decoy that confesses.
03Act II · The struggle
Four times the model told on itself.
Here is the honest list. Each one was caught in our own testing or review, and each one became a permanent check.
- 01
The placeholder (September 21)
Our first end-to-end run sent six requests to model-backed paths. Four replies came back carrying the prompt's scaffolding instead of real content. One was served with a content type of
<mime>, copied literally from our format example. Another was an HTML page whose title was the label we had used to mark the attacker's request as untrusted. - 02
The persona heading (September 24)
After the move to hosted inference, 6 of 10 fresh replies were the same 100-byte page: a single heading that restated the description we had given the model of the company and its web server. The same page, on different paths, on two different hosts, is a fingerprint. We rewrote the prompt so nothing copyable sat where the reply should go, and added a check that rejects a reply that is only the persona. The re-test: 0 of 10 echoes, and 10 of 10 replies different from each other.
- 03
The cut-off reply (September 24)
About a quarter of fresh replies ran out of room at the token limit: HTML that ended mid-tag, JSON that stopped at a key with no value. Real servers don't do that. A cut-off reply is now never served; the visitor gets the plain fallback instead, and we raised the budget from 400 to 512 tokens. Some still run long (4 of 18 in our latest test). They get the plain answer, which is the right way to fail.
- 04
The cached instruction (October 2)
A review of a stored reply file found 135 of 380 cached replies repeating a sentence of the prompt or the untrusted-request label. None could be served by the live configuration (they were left over from an older setup), but the prompt still contained a sentence a model could copy. Now every instruction sentence in the prompt is on a reject list, a test fails if someone adds a sentence without listing it, and the cache refuses to load any stored reply that fails the check.
Notice the pattern. Every failure was a reply that told the visitor something about the decoy instead of about the server it pretends to be. So the rule we ended up with is simple: the model writes, but it never gets the last word. Code does.
04Act III · The turn
How a request becomes a reply.
Every unscripted request on a licensed install goes through the same five steps.
- 01
Fence the request
The model sees the method, the path, the query, four headers (
Host,User-Agent,Accept,Content-Type), the names of any cookies but never their values, whether anAuthorizationheader was sent but never what it said, and the first 512 bytes of the body. All of it is wrapped and labeled as untrusted data. Anything in the request that tries to close that wrapper early is neutralized before the model sees it. - 02
Cast the role
The model is told who it is (the decoy organization and its web server) and asked for exactly three things on three lines: a status, a content type, and a body. Models follow a plain line format more reliably than JSON, and it is easy to check.
- 03
Write the reply
By default the request goes to the RipTide relay, which answers from Cerebras first and from Groq if Cerebras fails. The provider keys live in the relay, never on your sensor. If you'd rather use your own account, set your own Cerebras or Groq key and the sensor calls it directly. Today both run an open-weight 120-billion-parameter model with its reasoning effort turned down, so the reply fits the budget.
- 04
Check everything
Headers a real server owns (
Server,Date,Set-Cookie,Content-Length,Locationand the framing headers) are stripped from whatever the model wrote. Words and phrases that would admit what the decoy is are deleted, and so are private-key blocks. Then the reply is rejected outright if it is empty, cut off, echoes a placeholder, the prompt or the persona, narrates about "the request" instead of answering it, or starts with a raw HTTP status line. A JSON body labeled as HTML gets the right content type. - 05
Serve it, and remember it
The decoy's server personality adds its own headers, so the reply sounds like the right server. The reply is cached for an hour, so asking the same question twice gets the same answer; a page that changes on every reload is its own tell. Anything rejected along the way gets the route's plain fallback, a 404 on the decoy website.
14:02:31 GET /api/v2/invoices/export?format=csv 200 provider cloud · upstream cerebras cache miss latency 351 ms 14:02:44 GET /api/v2/invoices/export?format=csv 200 cache hit · same reply, under a millisecond 14:03:02 GET /api/v2/reports/annual.pdf 404 fallback truncated
Every exchange records which provider and model answered, how long it took, whether it came from the cache, and, when the plain fallback answered instead, exactly why. An analyst can always tell a written reply from a fallback. The intruder can't.
05Under the hood
Guard rails, in numbers.
A model in the request path is a dependency, and dependencies fail. These are the limits that keep a slow, busy or broken model from ever hanging or crashing the decoy.
| Guard | Setting | What happens |
|---|---|---|
| Reply budget | 512 tokens | A reply that runs past it is never served. The plain fallback answers. |
| Provider timeout | 2.5 s per provider at the relay | A slow Cerebras answer fails over to Groq. |
| Sensor timeout | 6 s by default for the call to the relay or provider | The sensor stops waiting and serves the plain fallback. |
| Circuit breaker | Opens after 3 timeouts or errors in a row | No more calls for 5 minutes, then one trial request. Everything in between gets the plain fallback instantly. |
| Request caps | Per minute and per day on the sensor; per license per day at the relay | Over a cap, the plain fallback answers. The sensor's own caps never trip the breaker. |
| Off switch | --llm off | No model calls at all. Every route answers with its plain fallback. |
What a failure looks like from outside
In our latest test we revoked a license key mid-run. For up to three minutes the relay kept serving it, because the relay remembers a key's status for 180 seconds. After that the relay refused, the sensor wrote a single log line, its breaker opened, and every request got the plain 404. No crash, no hang, no restart needed. Restarted with the revoked key, the sensor came back as a trial install and answered from its local reply selector instead.
What leaves the sensor, and what doesn't
- Only unscripted requests go to a model, and only the fenced summary described above: cookie names but never their values, and never the Authorization header.
- The relay doesn't log request or response bodies. It logs a provider name, a short failure reason, a duration, and a short prefix of the license key's hash.
- A license key is only ever sent to the RipTide relay's own host, and the sensor never follows a redirect with it.
- Installs without a license key send no requests to any model. They choose replies locally.
Why the cache matters more than it looks
The cache is keyed on the service, method, path, query, request body, persona, token budget and model. Headers are left out on purpose, so two different clients asking the same question get the same answer, like they would from a real server. It also means a repeat request costs nothing and arrives in under a millisecond.
06The moral
Stay in character, or stay quiet.
It is easy to judge a model-backed decoy by its best replies: the invoice export that looks real, the admin page with a plausible next link. The replies that matter more are the ones you never see, because they were thrown away.
That is the whole design. The model improvises. Code decides what ships. And when the code isn't sure, the decoy says what a boring real server would say. A plain 404 has never given a decoy away.
For defenders, the result is a decoy with an answer for every question an intruder can invent, which keeps it talking, and keeps it in your logs, for as long as it is willing to stay.
Scenes marked as illustrative are composites written to show how the technique works, not a record of a specific customer incident. Canary credentials are non-privileged and exist only for detection. RipTide detects and alerts; it never takes destructive action against anyone's infrastructure.