AI infrastructure

Guard the new crown jewels: your AI stack.

Model servers, MCP tools and GPU boxes are what attackers hunt for now: free compute, private data, and a foothold that speaks fluent API. RipTide plants decoy Ollama, vLLM and MCP endpoints that look just like the ones leaking across the internet.

The problem

AI infrastructure went up fast. Security came later.

Teams stood up model servers and MCP tools to move quickly, and plenty of them ship with no authentication at all. Attackers scan for them the way they once scanned for open databases: to steal compute, siphon data, or pivot deeper.

Running up someone else's model bill has a name now, LLMjacking, and the invoice lands on the victim. Your existing tools weren't built to tell an attacker's prompt from a developer's.

$100K+

a day in AI model charges a hijacked cloud account can run up, at full tilt

Sysdig Threat Research estimate, 2024

175K

Ollama hosts seen exposed to the internet

SentinelLABS & Censys, 2026

492

internet-exposed MCP servers running with no authentication at all

Trend Micro, July 2025

How RipTide catches it

Decoy model servers, with nothing behind them.

They look like the exposed model servers attackers already scan for. Nothing real sits behind them.

  1. 01

    Expose

    Decoy Ollama and vLLM endpoints answer like unauthenticated model servers, just like the ones that leak in the wild.

  2. 02

    Answer as MCP

    A decoy MCP server answers initialize and tools/list in both protocol versions, and lists tools no legitimate user would call.

  3. 03

    Record

    Every request is captured, prompts included, so you see what the attacker wanted the model to do.

  4. 04

    Investigate

    Nobody on your team uses the decoys, so any touch is a finding, with the full exchange one click away.

Detection coverage

Detection across AI infrastructure attacks.

From the first scan to the first tool call.

Attacks on AI infrastructure, the RipTide decoy that answers each, and what you learn
What the attacker doesThe decoy that answersWhat you learn
Scans for open model serversOllama & vLLM decoysWho is hunting for your AI stack, and from where.
Lists models and sends promptsOllama & vLLM decoysThe prompts they wanted run, captured word for word.
Connects to an MCP server and lists its toolsMCP server decoyWho connected to a tool server nobody configured, and which tools it saw.
Calls a tool that promises secretsMCP tools like read_vault_secretWhat it reached for, and the exact arguments it sent.
Probes for admin and inference APIsAn answer for any pathA believable answer to any path, and every request recorded.

The detections that do the work

The servers they're scanning for.

Every decoy has zero legitimate users. So every touch is a finding.

See every detection

Ollama & vLLM decoys

AI infra

Model servers that answer like unauthenticated Ollama and vLLM, run no model at all, and record every prompt.

MCP server decoy

AI agents

An internal-looking tool server that speaks both versions of MCP, lists tools like read_vault_secret, and asks every caller to introduce itself.

An answer for any path

Every plan

Probes for WordPress, .git, Jenkins, cloud metadata or Kubernetes get a believable reply instead of a 404: chosen on the box in Trial, written by an LLM in Professional.

What lands in your SOC

See exactly what they wanted your models to do.

A decoy model server records the whole conversation, so you learn what the attacker was after, not just that they knocked.

  • Every prompt and tool call, captured
  • Related sessions linked only when independent kinds of evidence agree
  • OCSF and STIX exports, and a TAXII feed your tools can pull

Illustrative example. Canary credentials are non-privileged and exist only for detection. RipTide detects and alerts; it never takes destructive action against anyone's infrastructure.

Questions

Fair questions.

What is LLMjacking?

Using stolen credentials or exposed servers to run someone else's AI models, leaving the owner with the bill and sometimes with their data exposed. Sysdig's threat research team coined the term in 2024.

Do the decoys run a real model?

No. They answer the way real Ollama and vLLM servers do, down to the error messages, and run no model at all. Requests the decoys have no script for still get a believable reply: chosen on the box in Trial, written by an LLM in Professional, under per-minute and daily caps.

Could the decoys expose our real models?

No. The decoys are separate services on the sensor with nothing real behind them.

Make contact.

Tell us what your AI stack looks like. We'll show you which decoys belong beside it.

Book a briefing

Thirty minutes with the people who built it. Bring your hardest question.

  • Watch a live agent set off a detection
  • Map decoys to your crown jewels
  • Plan a first deployment in one sitting

We use your email only to reply. No newsletter, no list, no sharing.

Keyboard shortcuts

T
Change the theme. Shift+T goes back.
?
Show this list
Esc
Close whatever's open