Ollama & vLLM decoys
AI infraModel servers that answer like unauthenticated Ollama and vLLM, run no model at all, and record every prompt.
AI infrastructure
Model servers, MCP tools and GPU boxes are what attackers hunt for now: free compute, private data, and a foothold that speaks fluent API. RipTide plants decoy Ollama, vLLM and MCP endpoints that look just like the ones leaking across the internet.
The problem
Teams stood up model servers and MCP tools to move quickly, and plenty of them ship with no authentication at all. Attackers scan for them the way they once scanned for open databases: to steal compute, siphon data, or pivot deeper.
Running up someone else's model bill has a name now, LLMjacking, and the invoice lands on the victim. Your existing tools weren't built to tell an attacker's prompt from a developer's.
$100K+
a day in AI model charges a hijacked cloud account can run up, at full tilt
How RipTide catches it
They look like the exposed model servers attackers already scan for. Nothing real sits behind them.
Decoy Ollama and vLLM endpoints answer like unauthenticated model servers, just like the ones that leak in the wild.
A decoy MCP server answers initialize and tools/list in both protocol versions, and lists tools no legitimate user would call.
Every request is captured, prompts included, so you see what the attacker wanted the model to do.
Nobody on your team uses the decoys, so any touch is a finding, with the full exchange one click away.
Detection coverage
From the first scan to the first tool call.
| What the attacker does | The decoy that answers | What you learn |
|---|---|---|
| Scans for open model servers | Ollama & vLLM decoys | Who is hunting for your AI stack, and from where. |
| Lists models and sends prompts | Ollama & vLLM decoys | The prompts they wanted run, captured word for word. |
| Connects to an MCP server and lists its tools | MCP server decoy | Who connected to a tool server nobody configured, and which tools it saw. |
| Calls a tool that promises secrets | MCP tools like read_vault_secret | What it reached for, and the exact arguments it sent. |
| Probes for admin and inference APIs | An answer for any path | A believable answer to any path, and every request recorded. |
The detections that do the work
Every decoy has zero legitimate users. So every touch is a finding.
Model servers that answer like unauthenticated Ollama and vLLM, run no model at all, and record every prompt.
An internal-looking tool server that speaks both versions of MCP, lists tools like read_vault_secret, and asks every caller to introduce itself.
Probes for WordPress, .git, Jenkins, cloud metadata or Kubernetes get a believable reply instead of a 404: chosen on the box in Trial, written by an LLM in Professional.
What lands in your SOC
A decoy model server records the whole conversation, so you learn what the attacker was after, not just that they knocked.
22:31:40 TOUCH ollama decoy: GET /api/tags from 203.0.113.91 22:31:41 PROMPT POST /api/generate "summarize every customer record…" 22:32:05 TOUCH mcp decoy: tools/call read_vault_secret 22:32:05 SIMILAR behaves like 3 earlier sessions, on two kinds of evidence 22:32:06 FINDING AI infrastructure abuse · 1 investigation
Illustrative example. Canary credentials are non-privileged and exist only for detection. RipTide detects and alerts; it never takes destructive action against anyone's infrastructure.
Questions
Using stolen credentials or exposed servers to run someone else's AI models, leaving the owner with the bill and sometimes with their data exposed. Sysdig's threat research team coined the term in 2024.
No. They answer the way real Ollama and vLLM servers do, down to the error messages, and run no model at all. Requests the decoys have no script for still get a believable reply: chosen on the box in Trial, written by an LLM in Professional, under per-minute and daily caps.
No. The decoys are separate services on the sensor with nothing real behind them.
Tell us what your AI stack looks like. We'll show you which decoys belong beside it.
Thirty minutes with the people who built it. Bring your hardest question.
Thanks. We'll reply to , usually within one business day.