Prefix caching makes long prompts cheap: when a request starts with tokens the engine has already processed, it reuses their KV blocks and skips most of the prefill. By default that reuse is global. Every request that begins the same way hits the same blocks, whoever sent it. That is a feature for a single team and a timing side channel for a service shared by several customers. Here is how the channel works, what vLLM provides against it, and what the API front end in front of vLLM has to do.
A cache hit is visible in the time to first token
vLLM splits a prompt into fixed-size blocks and gives each block a hash that covers the tokens of the block and the hash of the block before it. Two prompts with the same first N blocks therefore have the same first N hashes, and the second prompt finds the first one's blocks in the cache. Nothing in the hash says who sent the tokens. The same holds for an external KV tier behind vLLM, as long as it keys its objects by the same hashes.
A hit skips the prefill of the cached part, and the prefill is most of the time to first token (TTFT) on a long prompt. For example, on GLM-5.3 a prompt of about 3,000 tokens answers its first token in 672 ms from scratch and in 92 ms when its blocks are already in the cache. That is a seven-fold gap, larger than any network jitter, and every client can measure it.
Guess, time, repeat
An attacker who shares the deployment with a victim does not need to read the victim's cache. It is enough to ask whether a candidate text is in it. The attacker sends a guess and measures the TTFT: a fast answer means someone has already sent exactly those tokens. Guesses can be as narrow as the next block of a known template, such as the diagnosis field of a medical form, the name in a letter or the next lines of a competitor's system prompt. Extended block by block, they reconstruct the victim's input. This is not a theoretical concern: several papers demonstrated it on open-source serving engines in 2024–2025, and in vLLM it was assigned CVE-2025-46570. A usage counter with the number of cached prompt tokens, where the API returns one, gives the same answer without any timing.
A salt per request, nothing more
Since version 0.9, vLLM accepts a cache_salt field in the request body: chat
completions, completions, the Responses API, pooling endpoints and, in current versions, the
Anthropic-compatible /v1/messages. The salt is mixed into the hash of the first block
of the prompt. Every later block hashes its parent, so the salt changes every hash of the request,
including the blocks written during decoding and those of sliding-window and Mamba cache groups.
Requests with different salts never share a block, and requests with the same salt share exactly as
before.
{
"model": "your-model",
"messages": [{"role": "user", "content": "..."}],
"cache_salt": "<set by the front end, never by the client>"
}
Around the salt, vLLM hashes blocks with SHA-256 by default, so hashes cannot be forced to collide,
and reports the number of cached prompt tokens in usage only when started with
--enable-prompt-tokens-details. The Responses API reports it regardless of that flag.
The vLLM security
guide recommends a salt on every request in multi-tenant deployments and asks to treat the salt as
a secret.
That is where vLLM stops. It does not authenticate anyone, does not choose the salt, accepts whatever salt the client sends and treats an empty salt as none. The seed of the hash chain is a public constant unless the operator sets one, so without a salt the block hashes of a text can be computed by anyone who knows the text. Isolation is only as good as the component that sets the salt, and that component is the API front end.
Six rules for the gateway in front of vLLM
- Set the salt on every request, from the authenticated identity. Drop any
cache_saltthe client sent. Also drop the fields that reach the engine or its connectors unfiltered, such askv_transfer_paramsandvllm_xargs, and client-supplied multimodaluuidvalues, which feed the same hash. - Make the salt unguessable. Derive it with HMAC-SHA256 from a secret of at least 32 bytes that only the front end holds. A salt equal to the tenant name or account id can be guessed, and a guessed salt reopens the channel.
- Pick the scope to match the trust boundary. Per tenant shares the cache between colleagues, per user shares nothing between users, per session or request disables reuse across requests altogether. A client may narrow its scope, never widen it.
- Expose only the endpoints that carry the salt. vLLM must not be reachable around the front end, and endpoints without a salt field stay closed.
- Close the internal surfaces.
/metricsand/loadare not covered by the API key and count hits globally. The KV events stream carries tokens and salts in the clear. The storage ports of an external KV tier must not be reachable from tenant networks. - Keep the salt stable. Derived from the secret and the identity, the same user gets the same salt after a restart, and an external KV cache keeps serving that scope. Rotating the secret empties the cache for everyone, which is the right tool after a leak and the wrong one on a schedule.
A minimal version of this addition, and where it goes, is in the summary at the end of this post.
The same attack, with and without the salt
We ran the attack against our own deployment with the ADP KV cache connector, where the KV cache also lives on the accelerator behind vLLM. A victim tenant sent a prompt; the attacker then probed that prompt and, as a control, a fresh prompt of the same length that nobody had sent. The GPU prefix cache was reset between probes, so any hit had to come from the external tier. Thirty probes per group.
Without a salt, the probes of the victim's prompt and of the fresh prompt do not overlap at all: the slowest hit, 127 ms, is five times faster than the fastest miss. A single probe is enough to tell them apart (Kolmogorov–Smirnov test, p = 1.8 × 10⁻¹⁴). With a per-tenant salt, the two groups have medians of 676 and 679 ms and the same spread (p = 0.997): the attacker's probe of the victim's prompt takes exactly the path of any new prompt. Inside the victim's tenant nothing changed. A second user of that tenant got 2,944 cached tokens, the same as without isolation, and after a restart of vLLM the accelerator still served that tenant and no other.
external KV cache
The salt has to reach the storage key. An offload layer that keys its objects by vLLM's block hashes inherits the salt for free. A layer that hashes token ids itself drops it and moves the leak from the GPU to the storage tier: the attacker then gets hits from storage instead of from GPU memory. The ADP connector takes its keys only from vLLM's salted block hashes, and on shared storage can add a namespace per deployment.
This is how the rest of the industry does it
None of the systems below closes this channel by hiding the timing: a hit that is slowed down to look like a miss is a cache that no longer saves anything. They all close it by the same move. A trusted component puts a scope key at the root of the hash chain, and every cache tier below keys its data by those hashes. The engines differ only in the name of the field and in who is trusted to set it.
| System | Scope key | Who sets it |
|---|---|---|
| vLLM | cache_salt, mixed into the first block hash |
the caller; the security guide says the front end, per tenant |
| SGLang | cache_salt and extra_key in the radix-tree key and at the root of the storage-tier hash chain |
the caller |
| NVIDIA Dynamo | salt in the router's hash seed; a trusted x-tenant-id header overrides the salt in the body |
the gateway, which "must still authenticate the tenant identity" |
| vLLM offloading connector, Mooncake store | storage keys built from vLLM's salted block hashes; Mooncake adds a deployment prefix | inherited from vLLM |
| ADP KV cache connector | storage keys built only from vLLM's salted block hashes, plus an optional deployment namespace | inherited from vLLM; verified on hardware |
| Hosted APIs: OpenAI, Anthropic, AWS Bedrock, Google Vertex AI | the cache is scoped to the organization, workspace, account or project | the provider |
The hosted APIs say it in their own words. OpenAI: "Caches are not shared across organizations", and inside one organization a per-user cache key "help[s] prevent cache-hit probing across users". Anthropic: "Different organizations never share caches, even if they use identical prompts", with caches isolated per workspace on the Claude API. And vLLM's maintainers treat a cache tier or connector that loses the salt as a recurring class of bugs, with a conformance suite proposed to catch it.
The ADP KV cache connector follows the same model as vLLM's own offloading connector and the Mooncake store. It does not hash tokens itself: every key it writes to the accelerator comes from vLLM's salted block hashes, in every cache group, including hybrid Mamba layers, MTP and the sparse-attention indexer. On shared storage it can add a namespace per deployment, like Mooncake's key prefix. The measurement above is the proof: no hit crosses tenants, the attacker's timing carries no signal, and inside a tenant the cache works exactly as before. In cache isolation it is as strong as vLLM itself.
Isolation has a price. Pay it where the boundary is
A salt splits the cache. A prefix common to two scopes, such as a long shared system prompt or the same document, is computed and stored once per scope instead of once overall. Inside a scope nothing changes. So the question is where the trust boundary runs, not whether isolation is good.
| Deployment | Salt | Why |
|---|---|---|
| One team or one application, one trust domain | none | Nobody to hide prompts from; a salt only costs hits. |
| Shared public system prompt, few-shot examples, documentation assistant | none or per tenant | The shared prefix is the point of the cache; there is nothing secret in it. |
| API or SaaS serving several customers from one deployment | per tenant, at least | Customers must not learn each other's prompts; colleagues still share. |
| Users of one organization with private data: HR, medical, legal, personal mail | per user | The boundary runs between people, not between companies. |
| One-off prompts with secrets, such as credentials or keys pasted into a chat | per session or request | No reuse at all. Use it narrowly: it disables prefix caching for those requests. |
The limits of a salt
- Inside a scope everything is shared by design, including the cached-token counters. Make the scope no wider than the trust boundary.
- Load and eviction. Tenants still share GPU memory and storage capacity. A tenant can notice that the system is busy or flush others' blocks out of the cache, but cannot learn what their prompts contain.
- Other channels. Speculative decoding leaks through the number of tokens accepted per step, and streaming leaks through packet sizes and timing. A cache-aware scheduler that serves longer hits first leaks through queue order; vLLM schedules first-come-first-served or by priority. These need their own measures.
- The operator. Whoever runs the host, holds the front end's secret or reads the KV events stream is outside this model.
If you need isolation, do it in the gateway
Isolating the KV cache between tenants takes no change in vLLM and none in the KV cache tier behind it, as long as that tier keys its data by vLLM's block hashes. The work is in the gateway that already authenticates your clients:
- derive a salt per scope with HMAC under a secret only the gateway holds, and set it as
cache_salton every request; - drop any salt the client sent, together with
kv_transfer_paramsandvllm_xargs; - let clients reach vLLM only through the gateway, and keep
/metrics, KV events and the storage ports on the internal network; - pick the scope by the trust boundary: per tenant, per user or per session.
A minimal gateway in Python with FastAPI, for non-streaming chat requests. The same few lines go into any proxy you already run.
import base64, hashlib, hmac
import httpx
from fastapi import FastAPI, Request
SECRET = load_secret() # at least 32 random bytes, known only to the gateway
app = FastAPI()
vllm = httpx.AsyncClient(base_url="http://vllm:8000", timeout=600)
def cache_salt(scope_id: str) -> str:
digest = hmac.new(SECRET, f"v1|prod|tenant|{scope_id}".encode(), hashlib.sha256).digest()
return base64.urlsafe_b64encode(digest).rstrip(b"=").decode()
@app.post("/v1/chat/completions")
async def chat(request: Request):
tenant = authenticate(request) # your existing auth: API key -> tenant id
body = await request.json()
for field in ("cache_salt", "kv_transfer_params", "vllm_xargs"):
body.pop(field, None) # never trust the client with these
body["cache_salt"] = cache_salt(tenant) # the same tenant always gets the same salt
response = await vllm.post("/v1/chat/completions", json=body)
return response.json()
This is the arrangement the others recommend as well: the vLLM security guide asks for a secret salt on every request in multi-tenant deployments, NVIDIA Dynamo takes the tenant from a header its gateway has authenticated, and the hosted APIs draw the same line at the organization or workspace. With the ADP KV cache connector the isolation then extends to the accelerator by itself.
Sources
- vLLM advisory GHSA-4qjh-9fv9-r85r (CVE-2025-46570) and the fix, cache salting, vLLM PR #17045.
- vLLM security guide: cache salting and salt strategy.
- Song et al., The Early Bird Catches the Leak: Unveiling Timing Side Channels in LLM Serving Systems.
- Zheng et al., InputSnatch: Stealing Input in LLM Services via Timing Side-Channel Attacks.
- Wu et al., I Know What You Asked: Prompt Leakage via KV-Cache Sharing in Multi-Tenant LLM Serving, NDSS 2025.
- Prompt caching and cache isolation in hosted APIs: OpenAI, Anthropic.
- vLLM RFC #53194: a conformance suite for KV-cache key partitioning; issue #53495 is one instance, a connector keyed on raw tokens.
- Gu et al., Auditing Prompt Caching in Language Model APIs, ICML 2025.
Test setup
- Deployment
- GLM-5.3 FP8 on 8× NVIDIA H200 NVL, TP=8, vLLM 0.28.0 with prefix caching and the ADP KV cache connector; the KV cache is also stored on the accelerator reached over NVMe-oF.
- Front end
- A reference proxy that drops client-supplied salts and sets
cache_salt = HMAC-SHA256(K, "v1|deployment|tenant|id"); the key is generated at start and stored nowhere. - Timing test
- A prompt of about 3,000 tokens sent by the victim tenant; 30 attacker probes of that prompt and 30 of a fresh prompt of the same length, the GPU prefix cache reset between probes; two-sample Kolmogorov–Smirnov test. The control repeats the run without salts.
Running several tenants on one deployment?
The front-end rules above are short, and they decide whether a shared KV cache leaks. If you want to check your setup against the attack, we are happy to help.