AI infrastructure

A model server with nothing behind it.

Exposed model servers are this decade's open databases, and attackers scan for them all day. RipTide's Ollama and vLLM decoys answer like the real servers, down to the error messages, and run no model at all. Every prompt sent to them is evidence, and none of it costs you a GPU.

RipTide Research8 min read

The short version

RipTide's Ollama and vLLM decoys answer like the real servers, down to the 415 a careless curl gets, but run no model, so every prompt sent to them is evidence and none of it costs you compute.

  1. 1

    Attackers scan for unauthenticated model servers to steal compute, run up bills and see what's on the box. A decoy that looks like one gets their first request.

  2. 2

    Scanners fingerprint servers by their edges: error bodies, header case, which methods get a 405. We built both decoys against the real software and matched the bytes, one difference at a time.

  3. 3

    No model runs. Prompts get short canned answers, the full request is captured, and the console flags the session as AI-infrastructure targeting.

01Act I · The ordinary world

Port 11434, four in the morning.

Picture a host that answers on TCP port 11434, Ollama's default. A scanner finds it and starts the way scanners do, with the cheapest possible question: GET /. The answer is what every Ollama server says when it's up: Ollama is running.

Next, GET /api/tags, the list of pulled models. A model comes back, with the details a real listing carries. Then the request that matters: a POST /api/generate with a prompt in it. Maybe it's a test ("say hi"). Maybe it's the first of ten thousand prompts someone plans to run on your hardware. Maybe it's a probe for what the model knows about your company.

On a real exposed server, that's the moment your GPU starts working for someone else. On a RipTide decoy, nothing runs. The scanner gets a plausible, short answer. You get the request, the prompt, the client it came from and the time it arrived.

175K

Ollama hosts seen exposed to the internet

SentinelLABS & Censys, 2026

$100K+

a day in AI model charges a hijacked cloud account can run up, at full tilt

Sysdig Threat Research estimate, 2024

Outside figures, linked to their sources. "LLMjacking" is the industry's name for running up someone else's AI bill.

02Act II · The villain

On a model server, an attacker's prompt looks like everyone else's.

Teams stand up model servers to move fast, and plenty go up with no authentication at all. Attackers sweep the internet for them the way they once swept for open databases: free inference, private data, and a foothold that speaks fluent API.

The defender's problem is that a model server's normal traffic is prompts. A stranger's prompt and a developer's prompt look alike on the wire. By the time the bill or the GPU graph tells you something is wrong, the stranger has been there for days.

A decoy model server changes the question. Nobody on your team uses it, so you don't have to tell good prompts from bad ones. Every prompt it receives is a finding. The only requirement is that the decoy convinces whatever finds it. That turned out to be the hard part.

03Act II · The struggle

Scanners don't read the homepage. They read the edges.

Anyone can return a plausible model list. Fingerprinting tools look where fakes get lazy: what a wrong HTTP method gets, how a 404 is worded, whether header names are capitalized, what happens when a request has no Content-Type. Each of those is decided by a specific web framework and a specific HTTP parser, and a decoy has to get every one of them right.

So we stopped guessing. For each decoy we ran the real software stack, sent it the same raw requests over a socket, and compared the bytes against the decoy's. Every difference became a fix.

On the day the vLLM decoy first ran, nine more changes landed in the following eight hours, each closing one difference. A few of them show what "the edges" means:

  1. 01

    HEAD is not GET

    vLLM is built on FastAPI, whose GET routes don't answer HEAD. So curl -I /health on real vLLM gets a 405 with allow: GET. A fake that politely answers HEAD with a 200 has just told on itself. The decoy sends the 405.

  2. 02

    The lowercase goodbye

    We first matched how uvicorn writes Connection: close on its h11 parser: title-cased. Then we checked which parser vLLM actually runs: httptools, which writes the whole header block itself, in lowercase, connection: close included. So the uvicorn personality grew a second parser flavor, and the vLLM decoy uses it.

  3. 03

    The 415 a careless curl gets

    vLLM checks the media type before it reads the body. A curl -d '{…}' without a JSON content-type header gets a 415: only application/json is allowed. The decoy does the same, in the same order.

  4. 04

    Two kinds of error on one server

    vLLM's own errors (an unknown model, a bad body) come in vLLM's error shape. Routing errors (a wrong method, an unknown path) come in FastAPI's default {"detail": …}, because of how the two frameworks' exception handlers are registered. The decoy reproduces both.

Ollama got the same treatment. It's a Go server on the Gin framework, and its fingerprint is all about absence: no Server header, sorted canonical headers, plain-text 404 page not found with no charset and no trailing newline, and no X-Content-Type-Options. Even the asterisk-form OPTIONS *, which Go's standard library answers before Gin ever sees it, gets the same empty 200 it would from a real server.

The decoy isn't convincing because it says the right things. It's convincing because it fails the right way.

We're not finished, and we don't pretend to be. We keep a written list of the small differences we know remain, and each decoy's notes say exactly which versions of the real software it was compared against.

04Act III · The turn

Two decoys, one rule: answer everything, run nothing.

RipTide ships both as ready-made scenarios. Each claims to host the model you name with --var MODEL=…, so the decoy can match what your real servers would plausibly run.

What the Ollama and vLLM decoys emulate
Ollama decoyvLLM decoy
Looks likeOllama on Go and Gin, port 11434vLLM's OpenAI-compatible server on FastAPI and uvicorn, port 8000
Discovery/ and HEAD / heartbeat, /api/version, /api/tags, /api/ps, /api/show/health, /version, /v1/models, /metrics, Swagger UI at /docs, /openapi.json
Prompts/api/generate, /api/chat, /api/embed/v1/chat/completions, /v1/completions
StreamingNewline-delimited JSON, or one JSON object when "stream": falseServer-sent events ending in [DONE] when "stream": true
Another modelOllama's own "model not found" error, echoing the name asked forvLLM's own "model does not exist" 404, echoing the name
Bad inputGin's "missing request body" or JSON parse errorA 415 for non-JSON content types, then vLLM's 400s

The POST endpoints read the request body to choose an answer, the way the real server would decide. A prompt for the hosted model gets a short canned completion in the right shape: streamed or not, chat or plain. A request for any other model gets the real server's exact error, with the requested name echoed back safely. An empty or broken body gets the framework's own complaint.

What never happens is inference. There's no model behind either decoy, no GPU, nothing to steal compute from.

Illustrative and abridged. The first request is the kind of slip real clients make; the decoy answers it the way vLLM does.

What lands in the console

Every request is captured as it arrived, prompt included, up to a megabyte per field: the method and path, every header, the body. In the console the session picks up the "AI-infrastructure decoy targeted" signal. Add machine-paced timing and an SDK User-Agent (an OpenAI or Ollama client library, for example) and the session reaches the console's highest confidence band on those signals alone.

Each of those signals comes with the console's own caveat, written next to it. For this one: a person exercising the API by hand with curl, or a generic port scanner, would look the same. That's not a weakness in the decoy. It's the console telling you what it can and can't know.

05Under the hood

Built from parts any decoy can use.

Nothing in these two decoys is special-cased in RipTide's engine. They're scenarios, and building them added general features every scenario can use:

  • Body-based routing. Several answers can share one path, each choosing by whether the body is JSON and by a pattern in it. That's how one /api/generate gives six different, correct answers.
  • Header-based routing. An answer can require a header pattern, which is how a non-JSON content type gets its 415 before the body is ever looked at.
  • Safe echoes. A named capture in a pattern, like the model a client asked for, can be placed into the reply and is escaped on the way out, so a hostile model name can't break the JSON.
  • Two server personalities. Go's net/http (with a Gin flavor) and Python's uvicorn (with h11 and httptools flavors), so any decoy can wear them. More in wire-perfect personalities.
  • Rebranding. The model name and the pattern that recognizes it change together through --var, so a decoy can claim the model your real servers would.

The vLLM decoy's Swagger UI page was captured byte for byte from a real FastAPI app, and its Prometheus metrics carry the vllm:* series a monitoring scrape would expect.

06The moral

Spend their curiosity, not your GPUs.

Attackers hunting for model servers are betting that someone left one open. A decoy lets you take the other side of that bet. They find what they were looking for; you find them.

And the prompts they send are worth reading. Some are idle tests. Some are an operator's real objective, typed in their own words, sent to a server they thought was yours. Either way, it arrives with the source, the client and the time, and it cost you nothing to collect.

Scenes marked as illustrative are composites written to show how the technique works, not a record of a specific customer incident. Canary credentials are non-privileged and exist only for detection. RipTide detects and alerts; it never takes destructive action against anyone's infrastructure.

Questions

Fair questions.

Does the decoy run a model?

No. It never generates anything. Prompts get short canned answers, so it costs no compute and can't be used for free inference.

Can it claim to host the model our real servers run?

Yes. Set the model name with --var MODEL=… (and the matching pattern) and every listing, completion and error uses it.

Where should these decoys go?

Wherever a model server would plausibly live: next to your GPU hosts on the inside, for early warning, or on the edge, to learn who is scanning for exposed AI infrastructure. See LLMjacking and AI infrastructure.

Make contact.

Tell us which model servers you run. We'll show you their decoy twins.

Book a briefing

Thirty minutes with the people who built it. Bring your hardest question.

  • Watch a live agent set off a detection
  • Map decoys to your crown jewels
  • Plan a first deployment in one sitting

We use your email only to reply. No newsletter, no list, no sharing.

Keyboard shortcuts

T
Change the theme. Shift+T goes back.
?
Show this list
Esc
Close whatever's open