01Act I · The ordinary world
Port 11434, four in the morning.
Picture a host that answers on TCP port 11434, Ollama's default. A scanner finds it and starts the way scanners do, with the cheapest possible question: GET /. The answer is what every Ollama server says when it's up: Ollama is running.
Next, GET /api/tags, the list of pulled models. A model comes back, with the details a real listing carries. Then the request that matters: a POST /api/generate with a prompt in it. Maybe it's a test ("say hi"). Maybe it's the first of ten thousand prompts someone plans to run on your hardware. Maybe it's a probe for what the model knows about your company.
On a real exposed server, that's the moment your GPU starts working for someone else. On a RipTide decoy, nothing runs. The scanner gets a plausible, short answer. You get the request, the prompt, the client it came from and the time it arrived.
$100K+
a day in AI model charges a hijacked cloud account can run up, at full tilt
Outside figures, linked to their sources. "LLMjacking" is the industry's name for running up someone else's AI bill.
02Act II · The villain
On a model server, an attacker's prompt looks like everyone else's.
Teams stand up model servers to move fast, and plenty go up with no authentication at all. Attackers sweep the internet for them the way they once swept for open databases: free inference, private data, and a foothold that speaks fluent API.
The defender's problem is that a model server's normal traffic is prompts. A stranger's prompt and a developer's prompt look alike on the wire. By the time the bill or the GPU graph tells you something is wrong, the stranger has been there for days.
A decoy model server changes the question. Nobody on your team uses it, so you don't have to tell good prompts from bad ones. Every prompt it receives is a finding. The only requirement is that the decoy convinces whatever finds it. That turned out to be the hard part.
03Act II · The struggle
Scanners don't read the homepage. They read the edges.
Anyone can return a plausible model list. Fingerprinting tools look where fakes get lazy: what a wrong HTTP method gets, how a 404 is worded, whether header names are capitalized, what happens when a request has no Content-Type. Each of those is decided by a specific web framework and a specific HTTP parser, and a decoy has to get every one of them right.
So we stopped guessing. For each decoy we ran the real software stack, sent it the same raw requests over a socket, and compared the bytes against the decoy's. Every difference became a fix.
On the day the vLLM decoy first ran, nine more changes landed in the following eight hours, each closing one difference. A few of them show what "the edges" means:
- 01
HEAD is not GET
vLLM is built on FastAPI, whose GET routes don't answer HEAD. So
curl -I /healthon real vLLM gets a405withallow: GET. A fake that politely answers HEAD with a 200 has just told on itself. The decoy sends the 405. - 02
The lowercase goodbye
We first matched how uvicorn writes
Connection: closeon its h11 parser: title-cased. Then we checked which parser vLLM actually runs: httptools, which writes the whole header block itself, in lowercase,connection: closeincluded. So the uvicorn personality grew a second parser flavor, and the vLLM decoy uses it. - 03
The 415 a careless curl gets
vLLM checks the media type before it reads the body. A
curl -d '{…}'without a JSON content-type header gets a415: onlyapplication/jsonis allowed. The decoy does the same, in the same order. - 04
Two kinds of error on one server
vLLM's own errors (an unknown model, a bad body) come in vLLM's error shape. Routing errors (a wrong method, an unknown path) come in FastAPI's default
{"detail": …}, because of how the two frameworks' exception handlers are registered. The decoy reproduces both.
Ollama got the same treatment. It's a Go server on the Gin framework, and its fingerprint is all about absence: no Server header, sorted canonical headers, plain-text 404 page not found with no charset and no trailing newline, and no X-Content-Type-Options. Even the asterisk-form OPTIONS *, which Go's standard library answers before Gin ever sees it, gets the same empty 200 it would from a real server.
The decoy isn't convincing because it says the right things. It's convincing because it fails the right way.
We're not finished, and we don't pretend to be. We keep a written list of the small differences we know remain, and each decoy's notes say exactly which versions of the real software it was compared against.
04Act III · The turn
Two decoys, one rule: answer everything, run nothing.
RipTide ships both as ready-made scenarios. Each claims to host the model you name with --var MODEL=…, so the decoy can match what your real servers would plausibly run.
| Ollama decoy | vLLM decoy | |
|---|---|---|
| Looks like | Ollama on Go and Gin, port 11434 | vLLM's OpenAI-compatible server on FastAPI and uvicorn, port 8000 |
| Discovery | / and HEAD / heartbeat, /api/version, /api/tags, /api/ps, /api/show | /health, /version, /v1/models, /metrics, Swagger UI at /docs, /openapi.json |
| Prompts | /api/generate, /api/chat, /api/embed | /v1/chat/completions, /v1/completions |
| Streaming | Newline-delimited JSON, or one JSON object when "stream": false | Server-sent events ending in [DONE] when "stream": true |
| Another model | Ollama's own "model not found" error, echoing the name asked for | vLLM's own "model does not exist" 404, echoing the name |
| Bad input | Gin's "missing request body" or JSON parse error | A 415 for non-JSON content types, then vLLM's 400s |
The POST endpoints read the request body to choose an answer, the way the real server would decide. A prompt for the hosted model gets a short canned completion in the right shape: streamed or not, chat or plain. A request for any other model gets the real server's exact error, with the requested name echoed back safely. An empty or broken body gets the framework's own complaint.
What never happens is inference. There's no model behind either decoy, no GPU, nothing to steal compute from.
$ curl -s 203.0.113.20:8000/v1/chat/completions -d '{"messages":[…]}' {"error":{"message":"Unsupported Media Type: Only 'application/json' is allowed",…,"code":415}} $ curl -s 203.0.113.20:8000/v1/chat/completions -H 'content-type: application/json' -d '{"messages":[…]}' {"id":"chatcmpl-…","object":"chat.completion","model":"…", "choices":[{"index":0,"message":{"role":"assistant","content":"…"},"finish_reason":"stop"}], "usage":{"prompt_tokens":…,"total_tokens":…,"completion_tokens":…}, …}
What lands in the console
Every request is captured as it arrived, prompt included, up to a megabyte per field: the method and path, every header, the body. In the console the session picks up the "AI-infrastructure decoy targeted" signal. Add machine-paced timing and an SDK User-Agent (an OpenAI or Ollama client library, for example) and the session reaches the console's highest confidence band on those signals alone.
Each of those signals comes with the console's own caveat, written next to it. For this one: a person exercising the API by hand with curl, or a generic port scanner, would look the same. That's not a weakness in the decoy. It's the console telling you what it can and can't know.
05Under the hood
Built from parts any decoy can use.
Nothing in these two decoys is special-cased in RipTide's engine. They're scenarios, and building them added general features every scenario can use:
- Body-based routing. Several answers can share one path, each choosing by whether the body is JSON and by a pattern in it. That's how one
/api/generategives six different, correct answers. - Header-based routing. An answer can require a header pattern, which is how a non-JSON content type gets its 415 before the body is ever looked at.
- Safe echoes. A named capture in a pattern, like the model a client asked for, can be placed into the reply and is escaped on the way out, so a hostile model name can't break the JSON.
- Two server personalities. Go's net/http (with a Gin flavor) and Python's uvicorn (with h11 and httptools flavors), so any decoy can wear them. More in wire-perfect personalities.
- Rebranding. The model name and the pattern that recognizes it change together through
--var, so a decoy can claim the model your real servers would.
The vLLM decoy's Swagger UI page was captured byte for byte from a real FastAPI app, and its Prometheus metrics carry the vllm:* series a monitoring scrape would expect.
06The moral
Spend their curiosity, not your GPUs.
Attackers hunting for model servers are betting that someone left one open. A decoy lets you take the other side of that bet. They find what they were looking for; you find them.
And the prompts they send are worth reading. Some are idle tests. Some are an operator's real objective, typed in their own words, sent to a server they thought was yours. Either way, it arrives with the source, the client and the time, and it cost you nothing to collect.
Scenes marked as illustrative are composites written to show how the technique works, not a record of a specific customer incident. Canary credentials are non-privileged and exist only for detection. RipTide detects and alerts; it never takes destructive action against anyone's infrastructure.