What to monitor
| Signal | Where | What it tells you |
|---|---|---|
cached_tokens | usage.prompt_tokens_details of each API response (with --enable-prompt-tokens-details) | The hit of one request, GPU and card together. |
External prefix cache hit rate | vLLM’s periodic stats line | The share of looked-up tokens found on the card. |
vllm:external_prefix_cache_queries_total, vllm:external_prefix_cache_hits_total | vLLM /metrics | The same as counters; the hit rate is hits / queries over a window. vllm:prefix_cache_* is the GPU tier. |
External KV cache turned off on this rank | vLLM log | The gateway stopped answering; the external cache is off until the gateway and vLLM are restarted together. Page on it. |
Physical usage, CapacityUsage | pliocli system get_disk_usage, pliocli system get_status on the card’s host | How full the card is; alert well before the maximum (capacity). |
| Aggregated status | pliocli system get_status | Driver, service, RAID, firmware, temperature, supercapacitors, on-board flash. |
| New kinds of errors | gateway log | Anything other than the code: 135 misses below deserves a look. |
| Restarts | the gateway and vLLM containers | A gateway restart needs a vLLM restart too (Kubernetes). |
Read hit rates against what is possible: a prompt of N tokens can hit at most floor(N / chunk_size) × chunk_size
tokens, and the first request of every new prefix is a miss by definition. Compare the rate with and without the
connector on the same traffic rather than against 100 %.
No hits after a restart
A long prompt, sent before and after a restart of vLLM, comes back with cached_tokens: 0. Check, in this order:
PYTHONHASHSEEDis set, to the same value, in the vLLM environment. Without it every process hashes blocks differently.- The prompt is longer than one chunk, and identical, including the chat template and system prompt.
- Nothing changed in the stored layout since it was written: model and revision, tensor-parallel size,
chunk_size,key_namespace, gateway fragment settings (the list). - The gateway kept the cache: it starts with
--start-clean=0, and only vLLM was restarted. - The salt is stable, when the front end sets one: the same tenant must get the same
cache_saltafter the restart. Withrequire_cache_salt: truea request without a salt is never cached. - The card accepts writes:
CapacityUsage (OK)inpliocli system get_status, and physical usage growing while new prompts arrive. - The connector registered: the vLLM log has
XKV: N cache tensors registered for exchangeand noExternal KV cache turned off on this rank.
The hit rate falls while serving
| Cause | How it shows | What to do |
|---|---|---|
| The card is full and key eviction is off | CapacityUsage (CRITICAL), event 3002 in the storage service log; vLLM keeps serving, nothing new is cached | Enable key eviction; wipe the card if it must be emptied now. |
| The storage target restarted (remote card) | the NVMe device stops answering; the connector turns the external cache off | Reconnect the namespace, then restart the gateway and vLLM together. |
| The gateway went away | External KV cache turned off on this rank | Restart the gateway and vLLM together. |
| New traffic, new prefixes | the rate recovers as the cache fills | Nothing; size the card for the working set. |
Troubleshooting
Hybrid models stop with “Cannot get N free blocks”
ValueError: Cannot get 2 free blocks from the pool
Hybrid models (Mamba, Mamba-2, Gated DeltaNet, Kimi Delta Attention) need a small reserve of free KV blocks when vLLM
admits a request whose prefix comes from the card: each Mamba-type group takes one block for the restored state that
vLLM’s admission check does not count. Run them with --watermark=0.02, which keeps int(W × num_gpu_blocks) blocks
free at admission; the reserve must be at least the number of Mamba groups G:
int(W × num_gpu_blocks) ≥ G
0.02 covers three Mamba groups from a pool of 150 blocks up; on a smaller pool or with more groups take
W ≥ G / num_gpu_blocks (num_gpu_blocks is a label of vllm:cache_config_info on /metrics; the connector’s start
log has one group K: MambaSpec line per Mamba group). The reserve is about 2 % of the pool. Attention-only,
sliding-window, MLA and DSA models do not need it.
The gateway refuses vLLM: VERSION_MISMATCH
The gateway log shows SDK versions mismatch. Application version: <A>, SDK Version: <B> and the connector turns the
external cache off. The gateway and the vLLM image (or the kit it was built from) come from different releases: deploy
both from the same release.
Failed to open KVDatabase
vLLM and the gateway do not see the same /dev/shm: the gateway created the ring in its own. Run both with the host’s
IPC namespace (ipc: host, hostIPC: true) and remove any volume mounted over /dev/shm.
A hybrid model refuses to start
- Without
--enable-prefix-cachingvLLM turns the Mamba cache off, and the connector refuses to start. - The connector accepts only
--mamba-cache-mode align; Nemotron 3 Nano picks another mode by itself unless told. - Pipeline parallelism above 1 is refused for models with several groups per role.
A sliding-window model refuses to start
Deferred write of a sliding-window group (Gemma 4, GLM-4.7-Flash) needs synchronous scheduling and pipeline parallelism
1: add --no-async-scheduling, or set enable_deferred_write: false at some cost in decode speed. The start message
names the group.
The connector refuses to start, naming a cache
KV cache registry: 1 of 4 registered cache(s) have no writer the connector can reach; first of each kind: ...
The model allocates a cache the connector does not know how to exchange, and it refuses to start rather than silently caching nothing. Tell us the model and the line.
Out of memory on long prompts with GLM-5.x
Use --max-num-batched-tokens=8192: with 16384 the sparse MLA kernel runs out of memory on prompts from about 128k
tokens.
Long pauses on the first requests of DeepSeek-V4-Flash
DeepGEMM compiles its kernels on first use on vLLM 0.29, which pauses generation for 15–23 s a few times. Warm the
server up with representative requests and keep VLLM_CACHE_ROOT on a persistent volume so the compiled kernels
survive restarts.
The gateway log is full of NVMe media error type: 2 code: 135
Ordinary misses: the key is not on the card. Judge the card by its status and usage, not by the number of these lines (details).
Upgrades
- Read the release notes: a release that changes the stored layout says so, and then the card has to be wiped.
- Pull the new images by their versioned tags; never deploy a moving tag.
- Upgrade the gateway and the vLLM image together — the gateway accepts only an SDK of its own release. Stop vLLM, then the gateway; start the gateway, then vLLM.
- Verify a restore (Installation).
A new vLLM version or a different PYTHONHASHSEED can change the block hashes: the cache then refills from new
traffic, and nothing wrong is ever read.
Measuring the effect
- Measure under concurrency. Numbers from one request at a time are a lower bound and say little about a serving deployment: batching changes throughput by an order of magnitude more than any cache setting.
- Measure a restore after a restart. Hits from GPU memory and hits from the card look the same in
cached_tokens; restart vLLM (only vLLM) before the warm run to measure the card. - Alternate the setups (A, B, A, B) with at least three runs each, in quiet windows, and run the cold case only on a freshly started server: a second run in a row is already warm.
- Do not compare answers byte for byte. The engine does not reproduce its own outputs exactly, with or without a
cache — kernels are tuned anew at every start, and a cached prefix may take another numeric path than a recomputed
one. Judge a deployment by
cached_tokens, latency and the correctness of answers.
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.