Why it matters
Prefix caching shares computed KV blocks between every request that starts the same way, whoever sent it. A hit skips the prefill, and the prefill is most of the time to first token of a long prompt: on GLM-5.3 a prompt of about 3,000 tokens answers in 672 ms from scratch and in 92 ms from the cache. So any client can tell whether someone already sent a given text — guess, time, repeat — and reconstruct another tenant’s prompt block by block. In vLLM this is CVE-2025-46570. An external KV cache widens the window: the cache now outlives GPU memory and restarts.
For a single team or application there is nobody to hide prompts from and nothing to do. For a service shared by several customers, or by users with private data, isolate the cache as described here.
How isolation reaches the card
vLLM accepts a cache_salt field in the request body (chat completions, completions, the Responses API, pooling and
the Anthropic-compatible /v1/messages). The salt is mixed into the hash of the first block of the prompt, and every
later block hashes its parent, so the salt changes every block hash of the request. Requests with different salts never
share a block; requests with the same salt share exactly as before.
The ADP KV cache connector does not hash tokens itself. Every key it writes to the card is built from vLLM’s salted block hashes — in every cache group, including Mamba state, MTP layers and the sparse-attention indexer — so a tenant’s scope holds on the card exactly as it does in GPU memory, with no setting on the connector.
Isolation is only as good as the component that sets the salt. vLLM does not authenticate anyone, accepts whatever salt the client sends and treats an empty salt as none. That component is the API front end in front of vLLM.
1. Choose the scope
A salt splits the cache: a prefix common to two scopes — a long shared system prompt, the same document — is computed and stored once per scope instead of once overall. Inside a scope nothing changes. Put the boundary where trust ends:
| Deployment | Salt | Why |
|---|---|---|
| One team or one application, one trust domain | none | Nobody to hide prompts from; a salt only costs hits. |
| Shared public system prompt, few-shot examples, a documentation assistant | none or per tenant | The shared prefix is the point of the cache, and there is nothing secret in it. |
| An API or SaaS serving several customers from one deployment | per tenant, at least | Customers must not learn each other’s prompts; colleagues still share. |
| Users of one organization with private data: HR, medical, legal, personal mail | per user | The boundary runs between people, not between companies. |
| One-off prompts with secrets, such as credentials pasted into a chat | per session or request | No reuse at all — use it narrowly, or keep such requests out of the external cache with xkv_store: false. |
A client may narrow its scope, never widen it.
2. Set the salt in the front end
Six rules for the API front end:
- Set the salt on every request, from the authenticated identity. Drop any
cache_saltthe client sent, and the fields that reach the engine or its connectors unfiltered —kv_transfer_params,vllm_xargs, and client-supplied multimodaluuidvalues, which feed the same hash. - Make the salt unguessable. Derive it with HMAC-SHA256 from a secret of at least 32 bytes that only the front end holds. A salt equal to the tenant name or account id can be guessed, and a guessed salt reopens the channel.
- Match the scope to the trust boundary (above).
- Expose only the endpoints that carry the salt. vLLM must not be reachable around the front end, and endpoints without a salt field stay closed.
- Close the internal surfaces — next step.
- Keep the salt stable. Derived from the secret and the identity, the same tenant gets the same salt after a restart, and the ADP card keeps serving that scope. Rotating the secret empties the cache for everyone: the right tool after a leak, the wrong one on a schedule.
A front end in Python with FastAPI, for chat and completion requests, streaming included. The same few lines go into any proxy you already run.
import base64, hashlib, hmac, os
import httpx
from fastapi import FastAPI, HTTPException, Request
from fastapi.responses import JSONResponse, StreamingResponse
from starlette.background import BackgroundTask
SECRET = bytes.fromhex(os.environ["CACHE_SALT_SECRET"]) # 32+ random bytes, only the front end has them
DEPLOYMENT = "prod" # part of the salt: two deployments never share
STRIP = ("cache_salt", "kv_transfer_params", "vllm_xargs") # never taken from the client
app = FastAPI()
vllm = httpx.AsyncClient(base_url="http://vllm:8000", timeout=None)
def cache_salt(scope: str, scope_id: str) -> str:
"""The same scope always gets the same salt; nobody without SECRET can compute it."""
message = f"v1|{DEPLOYMENT}|{scope}|{scope_id}".encode()
digest = hmac.new(SECRET, message, hashlib.sha256).digest()
return base64.urlsafe_b64encode(digest).rstrip(b"=").decode()
def authenticate(request: Request) -> str:
"""Your existing authentication: API key -> tenant id."""
tenant = lookup_tenant(request.headers.get("authorization", ""))
if tenant is None:
raise HTTPException(status_code=401)
return tenant
def salted(body: dict, tenant: str) -> dict:
for field in STRIP:
body.pop(field, None)
for message in body.get("messages", []): # client-supplied multimodal ids feed the hash too
if isinstance(message.get("content"), list):
for part in message["content"]:
if isinstance(part, dict):
part.pop("uuid", None)
body["cache_salt"] = cache_salt("tenant", tenant) # or ("user", user_id) for a per-user scope
return body
async def forward(request: Request, path: str):
body = salted(await request.json(), authenticate(request))
if body.get("stream"):
upstream = await vllm.send(vllm.build_request("POST", path, json=body), stream=True)
return StreamingResponse(upstream.aiter_raw(), status_code=upstream.status_code,
media_type=upstream.headers.get("content-type"),
background=BackgroundTask(upstream.aclose))
upstream = await vllm.post(path, json=body)
return JSONResponse(upstream.json(), status_code=upstream.status_code)
@app.post("/v1/chat/completions")
async def chat_completions(request: Request):
return await forward(request, "/v1/chat/completions")
@app.post("/v1/completions")
async def completions(request: Request):
return await forward(request, "/v1/completions")
Generate the secret once (openssl rand -hex 32), keep it in your secret store, and give it to the front end only.
3. Close the internal surfaces
| Surface | Why | What to do |
|---|---|---|
| vLLM’s API port | Requests that bypass the front end carry no salt, or one the client chose. | Reachable from the front end only. |
/metrics, /load | Not covered by the API key; they count hits globally. | Internal network only. |
| KV events stream | Carries tokens and salts in the clear. | Internal network only, or off. |
| The gateway’s control port | Answers lookups and deletes without authentication. | --bind-address=127.0.0.1 when vLLM shares its network namespace; never on a tenant network. |
| The NVMe-oF fabric of a remote card | Storage traffic. | A storage network, not reachable from tenants. |
4. Turn on the connector’s switches
Isolation works without any connector setting. Two switches make it stricter, and one request field keeps a request out of the card entirely. All are off by default, and with them off the storage keys are byte for byte those of a deployment without isolation — so an upgrade needs no wipe.
| Setting | Use it when |
|---|---|
require_cache_salt: true | Your front end salts every request. A request it missed — no salt or an empty one — then neither reads nor stores the external cache (one warning on the first), instead of sharing its prompt with every other unsalted request. |
key_namespace: "<deployment>" | Several deployments share a card and could present the same block hashes: two models of the same geometry, two front ends with different salt policies. Up to 64 bytes. Setting or changing it starts a new cache. |
"kv_transfer_params": {"xkv_store": false} in a request | Request-scoped traffic whose salt nobody will match again: no lookup, no restore, no write; nothing derived from the prompt reaches the card. Set by the front end only. |
--kv-transfer-config '{"kv_connector":"FusIOnXConnector",
"kv_connector_module_path":"xkv_vllm_connector.fusionx_connector","kv_role":"kv_both",
"kv_connector_extra_config":{"xkv_gateway_port":5557,"chunk_size":128,
"require_cache_salt":true,"key_namespace":"prod-glm53"}}'
5. Verify it
Send the same long prompt through the front end as two tenants, before and after a restart of vLLM. Start vLLM with
--enable-prompt-tokens-details so responses report cached_tokens.
FRONT=http://front-end.internal:8080
PROMPT=$(python3 -c "print(' '.join(f'Line {i}: a private document of tenant A.' for i in range(3000)))")
ask() { # ask <api key>: cached prompt tokens of one request
curl -s "$FRONT/v1/completions" -H "Authorization: Bearer $1" -H 'Content-Type: application/json' \
-d "$(jq -n --arg p "$PROMPT" '{model: "glm-5.3", prompt: $p, max_tokens: 4}')" \
| jq '.usage.prompt_tokens_details.cached_tokens'
}
ask "$KEY_TENANT_A" # 0: the first time anyone sends it
ask "$KEY_TENANT_A" # close to the prompt length: same tenant, same salt
ask "$KEY_TENANT_B" # 0: another tenant, another key space
# restart vLLM (not the gateway), wait for /health, then:
ask "$KEY_TENANT_A" # close to the prompt length: restored from the ADP card
ask "$KEY_TENANT_B" # still 0
A non-zero answer for tenant B means a request reached vLLM without the front end’s salt, or with a salt shared between the tenants.
That is what we measured on GLM-5.3 with the ADP KV cache connector, with the GPU prefix cache reset between probes so that any hit had to come from the card: without a salt, the attacker’s probe of the victim’s prompt answered in 92 ms against 672 ms for a prompt nobody sent; with a per-tenant salt, 676 ms against 679 ms — no signal. A colleague in the victim’s tenant kept every cached token, and after a restart of vLLM the card still served that tenant and no other. The full study: Isolating the KV cache between tenants.
What a salt does not cover
- Inside a scope everything is shared by design, including the cached-token counters. Make the scope no wider than the trust boundary.
- Load and eviction. Tenants still share GPU memory and the card’s capacity. A tenant can notice that the system is busy or push others’ blocks out of the cache, but cannot learn what their prompts contain.
- Other channels. Speculative decoding leaks through the number of tokens accepted per step, streaming through packet sizes and timing, a cache-aware scheduler through queue order. These need their own measures.
- The operator. Whoever runs the hosts, holds the front end’s secret or reads the KV events stream is outside this model.
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.