Back to Blog
Security KV Cache vLLM

Isolating the KV cache between tenants: what vLLM does, what the front end must do

Awide Labs Engineering · September 28, 2026

Prefix caching makes long prompts cheap: when a request starts with tokens the engine has already processed, it reuses their KV blocks and skips most of the prefill. By default that reuse is global. Every request that begins the same way hits the same blocks, whoever sent it. That is a feature for a single team and a timing side channel for a service shared by several customers. Here is how the channel works, what vLLM provides against it, and what the API front end in front of vLLM has to do.

How the cache is shared

A cache hit is visible in the time to first token

vLLM splits a prompt into fixed-size blocks and gives each block a hash that covers the tokens of the block and the hash of the block before it. Two prompts with the same first N blocks therefore have the same first N hashes, and the second prompt finds the first one's blocks in the cache. Nothing in the hash says who sent the tokens. The same holds for an external KV tier behind vLLM, as long as it keys its objects by the same hashes.

A hit skips the prefill of the cached part, and the prefill is most of the time to first token (TTFT) on a long prompt. For example, on GLM-5.3 a prompt of about 3,000 tokens answers its first token in 672 ms from scratch and in 92 ms when its blocks are already in the cache. That is a seven-fold gap, larger than any network jitter, and every client can measure it.

The attack

Guess, time, repeat

An attacker who shares the deployment with a victim does not need to read the victim's cache. It is enough to ask whether a candidate text is in it. The attacker sends a guess and measures the TTFT: a fast answer means someone has already sent exactly those tokens. Guesses can be as narrow as the next block of a known template, such as the diagnosis field of a medical form, the name in a letter or the next lines of a competitor's system prompt. Extended block by block, they reconstruct the victim's input. This is not a theoretical concern: several papers demonstrated it on open-source serving engines in 2024–2025, and in vLLM it was assigned CVE-2025-46570. A usage counter with the number of cached prompt tokens, where the API returns one, gives the same answer without any timing.

One shared key space: the probe tells hit from miss
measured, GLM-5.3 FP8 · 8× H200 · vLLM 0.28
Tenant A · victim sends private prompt P Tenant B · attacker sends guesses, times them Front end authenticates, forwards without a salt vLLM prefix cache one key space for everyone hash = H(parent, block tokens) h₁(P) h₂(P) h₃(P) h₄(P) cached by A's request, found by anyone who sends P P G WHAT THE ATTACKER MEASURES guess G ≠ P · miss full prefill 672 ms guess G = P · hit prefill skipped 92 ms → G is what A sent TTFT, MEDIAN OF 30 PROBES · GLM-5.3 FP8, 8× H200, ~3,000-TOKEN PROMPT
What vLLM gives you

A salt per request, nothing more

Since version 0.9, vLLM accepts a cache_salt field in the request body: chat completions, completions, the Responses API, pooling endpoints and, in current versions, the Anthropic-compatible /v1/messages. The salt is mixed into the hash of the first block of the prompt. Every later block hashes its parent, so the salt changes every hash of the request, including the blocks written during decoding and those of sliding-window and Mamba cache groups. Requests with different salts never share a block, and requests with the same salt share exactly as before.

{
  "model": "your-model",
  "messages": [{"role": "user", "content": "..."}],
  "cache_salt": "<set by the front end, never by the client>"
}

Around the salt, vLLM hashes blocks with SHA-256 by default, so hashes cannot be forced to collide, and reports the number of cached prompt tokens in usage only when started with --enable-prompt-tokens-details. The Responses API reports it regardless of that flag. The vLLM security guide recommends a salt on every request in multi-tenant deployments and asks to treat the salt as a secret.

That is where vLLM stops. It does not authenticate anyone, does not choose the salt, accepts whatever salt the client sends and treats an empty salt as none. The seed of the hash chain is a public constant unless the operator sets one, so without a salt the block hashes of a text can be computed by anyone who knows the text. Isolation is only as good as the component that sets the salt, and that component is the API front end.

What the front end must do

Six rules for the gateway in front of vLLM

  1. Set the salt on every request, from the authenticated identity. Drop any cache_salt the client sent. Also drop the fields that reach the engine or its connectors unfiltered, such as kv_transfer_params and vllm_xargs, and client-supplied multimodal uuid values, which feed the same hash.
  2. Make the salt unguessable. Derive it with HMAC-SHA256 from a secret of at least 32 bytes that only the front end holds. A salt equal to the tenant name or account id can be guessed, and a guessed salt reopens the channel.
  3. Pick the scope to match the trust boundary. Per tenant shares the cache between colleagues, per user shares nothing between users, per session or request disables reuse across requests altogether. A client may narrow its scope, never widen it.
  4. Expose only the endpoints that carry the salt. vLLM must not be reachable around the front end, and endpoints without a salt field stay closed.
  5. Close the internal surfaces. /metrics and /load are not covered by the API key and count hits globally. The KV events stream carries tokens and salts in the clear. The storage ports of an external KV tier must not be reachable from tenant networks.
  6. Keep the salt stable. Derived from the secret and the identity, the same user gets the same salt after a restart, and an external KV cache keeps serving that scope. Rotating the secret empties the cache for everyone, which is the right tool after a leak and the wrong one on a schedule.

A minimal version of this addition, and where it goes, is in the summary at the end of this post.

Per-tenant salt: every probe looks the same
measured, same setup, salt set by the front end
Tenant A · victim sends prompt P Tenant A · colleague sends P again Tenant B · attacker sends guesses G Front end drops client salts, sets cache_salt from the tenant's identity s_A = HMAC(K, A) s_B = HMAC(K, B) vLLM prefix cache h₁ = H(salt, block tokens) KEY SPACE OF A h₁(s_A,P) h₂ h₃ KEY SPACE OF B h₁(s_B,G) h₂ h₃ hit miss WHAT THE ATTACKER MEASURES guess G ≠ P · miss full prefill 679 ms guess G = P · miss too other salt, other hashes 676 ms no signal TTFT, MEDIAN OF 30 PROBES · SAME SETUP, PER-TENANT SALT
What we measured

The same attack, with and without the salt

We ran the attack against our own deployment with the ADP KV cache connector, where the KV cache also lives on the accelerator behind vLLM. A victim tenant sent a prompt; the attacker then probed that prompt and, as a control, a fresh prompt of the same length that nobody had sent. The GPU prefix cache was reset between probes, so any hit had to come from the external tier. Thirty probes per group.

Attacker's time to first token, 30 probes per row
GLM-5.3 FP8 · 8× H200 NVL · TP=8 · vLLM 0.28 · ADP connector
0 200 400 600 800 the victim's cached prompt 86 ms 93 ms 100 ms 97 ms 88 ms 91 ms 96 ms 93 ms 92 ms 90 ms 127 ms 87 ms 95 ms 92 ms 86 ms 90 ms 90 ms 97 ms 87 ms 94 ms 72 ms 93 ms 95 ms 96 ms 83 ms 89 ms 90 ms 91 ms 103 ms 97 ms 92 ms a prompt nobody sent 668 ms 667 ms 679 ms 683 ms 680 ms 670 ms 669 ms 671 ms 673 ms 683 ms 676 ms 671 ms 676 ms 682 ms 670 ms 670 ms 674 ms 688 ms 676 ms 667 ms 679 ms 669 ms 678 ms 668 ms 669 ms 671 ms 681 ms 668 ms 683 ms 666 ms 672 ms the victim's cached prompt 673 ms 671 ms 674 ms 690 ms 669 ms 676 ms 674 ms 672 ms 671 ms 686 ms 670 ms 674 ms 676 ms 683 ms 672 ms 675 ms 672 ms 680 ms 668 ms 681 ms 683 ms 686 ms 688 ms 711 ms 726 ms 695 ms 712 ms 676 ms 681 ms 689 ms 676 ms a prompt nobody sent 685 ms 680 ms 672 ms 687 ms 675 ms 675 ms 687 ms 673 ms 668 ms 687 ms 679 ms 671 ms 670 ms 684 ms 672 ms 673 ms 671 ms 680 ms 671 ms 681 ms 673 ms 685 ms 680 ms 855 ms 699 ms 702 ms 693 ms 673 ms 674 ms 685 ms 679 ms NO ISOLATION PER-TENANT SALT TIME TO FIRST TOKEN, MS · 30 PROBES PER ROW · WHITE TICK: MEDIAN

Without a salt, the probes of the victim's prompt and of the fresh prompt do not overlap at all: the slowest hit, 127 ms, is five times faster than the fastest miss. A single probe is enough to tell them apart (Kolmogorov–Smirnov test, p = 1.8 × 10⁻¹⁴). With a per-tenant salt, the two groups have medians of 676 and 679 ms and the same spread (p = 0.997): the attacker's probe of the victim's prompt takes exactly the path of any new prompt. Inside the victim's tenant nothing changed. A second user of that tenant got 2,944 cached tokens, the same as without isolation, and after a restart of vLLM the accelerator still served that tenant and no other.

If you run an
external KV cache

The salt has to reach the storage key. An offload layer that keys its objects by vLLM's block hashes inherits the salt for free. A layer that hashes token ids itself drops it and moves the leak from the GPU to the storage tier: the attacker then gets hits from storage instead of from GPU memory. The ADP connector takes its keys only from vLLM's salted block hashes, and on shared storage can add a namespace per deployment.

Industry practice

This is how the rest of the industry does it

None of the systems below closes this channel by hiding the timing: a hit that is slowed down to look like a miss is a cache that no longer saves anything. They all close it by the same move. A trusted component puts a scope key at the root of the hash chain, and every cache tier below keys its data by those hashes. The engines differ only in the name of the field and in who is trusted to set it.

Where the scope key lives
public documentation and source code of each system
System Scope key Who sets it
vLLM cache_salt, mixed into the first block hash the caller; the security guide says the front end, per tenant
SGLang cache_salt and extra_key in the radix-tree key and at the root of the storage-tier hash chain the caller
NVIDIA Dynamo salt in the router's hash seed; a trusted x-tenant-id header overrides the salt in the body the gateway, which "must still authenticate the tenant identity"
vLLM offloading connector, Mooncake store storage keys built from vLLM's salted block hashes; Mooncake adds a deployment prefix inherited from vLLM
ADP KV cache connector storage keys built only from vLLM's salted block hashes, plus an optional deployment namespace inherited from vLLM; verified on hardware
Hosted APIs: OpenAI, Anthropic, AWS Bedrock, Google Vertex AI the cache is scoped to the organization, workspace, account or project the provider

The hosted APIs say it in their own words. OpenAI: "Caches are not shared across organizations", and inside one organization a per-user cache key "help[s] prevent cache-hit probing across users". Anthropic: "Different organizations never share caches, even if they use identical prompts", with caches isolated per workspace on the Claude API. And vLLM's maintainers treat a cache tier or connector that loses the salt as a recurring class of bugs, with a conformance suite proposed to catch it.

The ADP KV cache connector follows the same model as vLLM's own offloading connector and the Mooncake store. It does not hash tokens itself: every key it writes to the accelerator comes from vLLM's salted block hashes, in every cache group, including hybrid Mamba layers, MTP and the sparse-attention indexer. On shared storage it can add a namespace per deployment, like Mooncake's key prefix. The measurement above is the proof: no hit crosses tenants, the attacker's timing carries no signal, and inside a tenant the cache works exactly as before. In cache isolation it is as strong as vLLM itself.

When to turn it on

Isolation has a price. Pay it where the boundary is

A salt splits the cache. A prefix common to two scopes, such as a long shared system prompt or the same document, is computed and stored once per scope instead of once overall. Inside a scope nothing changes. So the question is where the trust boundary runs, not whether isolation is good.

Where to put the salt
scope = the widest group whose members may learn each other's prompts
Deployment Salt Why
One team or one application, one trust domain none Nobody to hide prompts from; a salt only costs hits.
Shared public system prompt, few-shot examples, documentation assistant none or per tenant The shared prefix is the point of the cache; there is nothing secret in it.
API or SaaS serving several customers from one deployment per tenant, at least Customers must not learn each other's prompts; colleagues still share.
Users of one organization with private data: HR, medical, legal, personal mail per user The boundary runs between people, not between companies.
One-off prompts with secrets, such as credentials or keys pasted into a chat per session or request No reuse at all. Use it narrowly: it disables prefix caching for those requests.
What it does not cover

The limits of a salt

  • Inside a scope everything is shared by design, including the cached-token counters. Make the scope no wider than the trust boundary.
  • Load and eviction. Tenants still share GPU memory and storage capacity. A tenant can notice that the system is busy or flush others' blocks out of the cache, but cannot learn what their prompts contain.
  • Other channels. Speculative decoding leaks through the number of tokens accepted per step, and streaming leaks through packet sizes and timing. A cache-aware scheduler that serves longer hits first leaks through queue order; vLLM schedules first-come-first-served or by priority. These need their own measures.
  • The operator. Whoever runs the host, holds the front end's secret or reads the KV events stream is outside this model.
Summary

If you need isolation, do it in the gateway

Isolating the KV cache between tenants takes no change in vLLM and none in the KV cache tier behind it, as long as that tier keys its data by vLLM's block hashes. The work is in the gateway that already authenticates your clients:

  • derive a salt per scope with HMAC under a secret only the gateway holds, and set it as cache_salt on every request;
  • drop any salt the client sent, together with kv_transfer_params and vllm_xargs;
  • let clients reach vLLM only through the gateway, and keep /metrics, KV events and the storage ports on the internal network;
  • pick the scope by the trust boundary: per tenant, per user or per session.
Where the change goes
one addition in the gateway; vLLM and the KV tier unchanged
Clients tenant A tenant B … API gateway auth: key → tenant + salt middleware drop client salt cache_salt = HMAC(K, tenant) vLLM salt enters every block hash KV cache tier connector keys = salted block hashes, as is no direct access to vLLM INTERNAL NETWORK ONLY /metrics · /load · KV events · storage ports of the KV tier THE ONLY NEW CODE IS THE PURPLE BOX · VLLM AND THE KV TIER STAY AS THEY ARE

A minimal gateway in Python with FastAPI, for non-streaming chat requests. The same few lines go into any proxy you already run.

import base64, hashlib, hmac
import httpx
from fastapi import FastAPI, Request

SECRET = load_secret()          # at least 32 random bytes, known only to the gateway
app = FastAPI()
vllm = httpx.AsyncClient(base_url="http://vllm:8000", timeout=600)

def cache_salt(scope_id: str) -> str:
    digest = hmac.new(SECRET, f"v1|prod|tenant|{scope_id}".encode(), hashlib.sha256).digest()
    return base64.urlsafe_b64encode(digest).rstrip(b"=").decode()

@app.post("/v1/chat/completions")
async def chat(request: Request):
    tenant = authenticate(request)              # your existing auth: API key -> tenant id
    body = await request.json()
    for field in ("cache_salt", "kv_transfer_params", "vllm_xargs"):
        body.pop(field, None)                   # never trust the client with these
    body["cache_salt"] = cache_salt(tenant)     # the same tenant always gets the same salt
    response = await vllm.post("/v1/chat/completions", json=body)
    return response.json()

This is the arrangement the others recommend as well: the vLLM security guide asks for a secret salt on every request in multi-tenant deployments, NVIDIA Dynamo takes the tenant from a header its gateway has authenticated, and the hosted APIs draw the same line at the organization or workspace. With the ADP KV cache connector the isolation then extends to the accelerator by itself.

Further reading

Sources

How we measured

Test setup

Deployment
GLM-5.3 FP8 on 8× NVIDIA H200 NVL, TP=8, vLLM 0.28.0 with prefix caching and the ADP KV cache connector; the KV cache is also stored on the accelerator reached over NVMe-oF.
Front end
A reference proxy that drops client-supplied salts and sets cache_salt = HMAC-SHA256(K, "v1|deployment|tenant|id"); the key is generated at start and stored nowhere.
Timing test
A prompt of about 3,000 tokens sent by the victim tenant; 30 attacker probes of that prompt and 30 of a fresh prompt of the same length, the GPU prefix cache reset between probes; two-sample Kolmogorov–Smirnov test. The control repeats the run without salts.

Running several tenants on one deployment?

The front-end rules above are short, and they decide whether a shared KV cache leaks. If you want to check your setup against the attack, we are happy to help.

Talk to Awide Labs