Documentation menu
Operations

Operations & troubleshooting

What to monitor, what to do when hits disappear, troubleshooting, upgrades, and how to measure the effect of the cache.

What to monitor

SignalWhereWhat it tells you
cached_tokensusage.prompt_tokens_details of each API response (with --enable-prompt-tokens-details)The hit of one request, GPU and card together.
External prefix cache hit ratevLLM’s periodic stats lineThe share of looked-up tokens found on the card.
vllm:external_prefix_cache_queries_total, vllm:external_prefix_cache_hits_totalvLLM /metricsThe same as counters; the hit rate is hits / queries over a window. vllm:prefix_cache_* is the GPU tier.
External KV cache turned off on this rankvLLM logThe gateway stopped answering; the external cache is off until the gateway and vLLM are restarted together. Page on it.
Physical usage, CapacityUsagepliocli system get_disk_usage, pliocli system get_status on the card’s hostHow full the card is; alert well before the maximum (capacity).
Aggregated statuspliocli system get_statusDriver, service, RAID, firmware, temperature, supercapacitors, on-board flash.
New kinds of errorsgateway logAnything other than the code: 135 misses below deserves a look.
Restartsthe gateway and vLLM containersA gateway restart needs a vLLM restart too (Kubernetes).

Read hit rates against what is possible: a prompt of N tokens can hit at most floor(N / chunk_size) × chunk_size tokens, and the first request of every new prefix is a miss by definition. Compare the rate with and without the connector on the same traffic rather than against 100 %.

No hits after a restart

A long prompt, sent before and after a restart of vLLM, comes back with cached_tokens: 0. Check, in this order:

  1. PYTHONHASHSEED is set, to the same value, in the vLLM environment. Without it every process hashes blocks differently.
  2. The prompt is longer than one chunk, and identical, including the chat template and system prompt.
  3. Nothing changed in the stored layout since it was written: model and revision, tensor-parallel size, chunk_size, key_namespace, gateway fragment settings (the list).
  4. The gateway kept the cache: it starts with --start-clean=0, and only vLLM was restarted.
  5. The salt is stable, when the front end sets one: the same tenant must get the same cache_salt after the restart. With require_cache_salt: true a request without a salt is never cached.
  6. The card accepts writes: CapacityUsage (OK) in pliocli system get_status, and physical usage growing while new prompts arrive.
  7. The connector registered: the vLLM log has XKV: N cache tensors registered for exchange and no External KV cache turned off on this rank.

The hit rate falls while serving

CauseHow it showsWhat to do
The card is full and key eviction is offCapacityUsage (CRITICAL), event 3002 in the storage service log; vLLM keeps serving, nothing new is cachedEnable key eviction; wipe the card if it must be emptied now.
The storage target restarted (remote card)the NVMe device stops answering; the connector turns the external cache offReconnect the namespace, then restart the gateway and vLLM together.
The gateway went awayExternal KV cache turned off on this rankRestart the gateway and vLLM together.
New traffic, new prefixesthe rate recovers as the cache fillsNothing; size the card for the working set.

Troubleshooting

Hybrid models stop with “Cannot get N free blocks”

ValueError: Cannot get 2 free blocks from the pool

Hybrid models (Mamba, Mamba-2, Gated DeltaNet, Kimi Delta Attention) need a small reserve of free KV blocks when vLLM admits a request whose prefix comes from the card: each Mamba-type group takes one block for the restored state that vLLM’s admission check does not count. Run them with --watermark=0.02, which keeps int(W × num_gpu_blocks) blocks free at admission; the reserve must be at least the number of Mamba groups G:

int(W × num_gpu_blocks) ≥ G

0.02 covers three Mamba groups from a pool of 150 blocks up; on a smaller pool or with more groups take W ≥ G / num_gpu_blocks (num_gpu_blocks is a label of vllm:cache_config_info on /metrics; the connector’s start log has one group K: MambaSpec line per Mamba group). The reserve is about 2 % of the pool. Attention-only, sliding-window, MLA and DSA models do not need it.

The gateway refuses vLLM: VERSION_MISMATCH

The gateway log shows SDK versions mismatch. Application version: <A>, SDK Version: <B> and the connector turns the external cache off. The gateway and the vLLM image (or the kit it was built from) come from different releases: deploy both from the same release.

Failed to open KVDatabase

vLLM and the gateway do not see the same /dev/shm: the gateway created the ring in its own. Run both with the host’s IPC namespace (ipc: host, hostIPC: true) and remove any volume mounted over /dev/shm.

A hybrid model refuses to start

  • Without --enable-prefix-caching vLLM turns the Mamba cache off, and the connector refuses to start.
  • The connector accepts only --mamba-cache-mode align; Nemotron 3 Nano picks another mode by itself unless told.
  • Pipeline parallelism above 1 is refused for models with several groups per role.

A sliding-window model refuses to start

Deferred write of a sliding-window group (Gemma 4, GLM-4.7-Flash) needs synchronous scheduling and pipeline parallelism 1: add --no-async-scheduling, or set enable_deferred_write: false at some cost in decode speed. The start message names the group.

The connector refuses to start, naming a cache

KV cache registry: 1 of 4 registered cache(s) have no writer the connector can reach; first of each kind: ...

The model allocates a cache the connector does not know how to exchange, and it refuses to start rather than silently caching nothing. Tell us the model and the line.

Out of memory on long prompts with GLM-5.x

Use --max-num-batched-tokens=8192: with 16384 the sparse MLA kernel runs out of memory on prompts from about 128k tokens.

Long pauses on the first requests of DeepSeek-V4-Flash

DeepGEMM compiles its kernels on first use on vLLM 0.29, which pauses generation for 15–23 s a few times. Warm the server up with representative requests and keep VLLM_CACHE_ROOT on a persistent volume so the compiled kernels survive restarts.

The gateway log is full of NVMe media error type: 2 code: 135

Ordinary misses: the key is not on the card. Judge the card by its status and usage, not by the number of these lines (details).

Upgrades

  1. Read the release notes: a release that changes the stored layout says so, and then the card has to be wiped.
  2. Pull the new images by their versioned tags; never deploy a moving tag.
  3. Upgrade the gateway and the vLLM image together — the gateway accepts only an SDK of its own release. Stop vLLM, then the gateway; start the gateway, then vLLM.
  4. Verify a restore (Installation).

A new vLLM version or a different PYTHONHASHSEED can change the block hashes: the cache then refills from new traffic, and nothing wrong is ever read.

Measuring the effect

  • Measure under concurrency. Numbers from one request at a time are a lower bound and say little about a serving deployment: batching changes throughput by an order of magnitude more than any cache setting.
  • Measure a restore after a restart. Hits from GPU memory and hits from the card look the same in cached_tokens; restart vLLM (only vLLM) before the warm run to measure the card.
  • Alternate the setups (A, B, A, B) with at least three runs each, in quiet windows, and run the cold case only on a freshly started server: a second run in a row is already warm.
  • Do not compare answers byte for byte. The engine does not reproduce its own outputs exactly, with or without a cache — kernels are tuned anew at every start, and a cached prefix may take another numeric path than a recomputed one. Judge a deployment by cached_tokens, latency and the correctness of answers.

Planning a deployment?

Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.

Talk to Awide Labs