Release v6.1.1 of the ADP KV cache connector is out. We finished it over the weekend: it carries per-tenant KV cache isolation all the way into the external store, loads model weights about six times faster, and no longer lets a lost storage gateway stall the engine. It passed a full regression on hardware and is already serving production traffic.
KV cache isolation between tenants
Prefix caching shares computed KV blocks between every request that starts the same way, and the time to first
token can tell anyone whether a prompt was already cached. vLLM closes that channel with a per-request
cache_salt. The connector now keys the external store only by vLLM's salted block hashes, so a
tenant's scope holds on the drives exactly as it does in GPU memory. Optional switches add a deployment
namespace, refuse to cache requests without a salt, and let the front end keep a single request out of the
cache entirely.
Everything is off by default, and storage keys stay byte-for-byte as before: an upgrade needs no cache wipe. Isolation switches on when the API front end sets the salt. In our end-to-end test, no request found another tenant's cache, and an attacker's timing probes could no longer tell a cached prompt from an uncached one. How it works and what the front end has to do: Isolating the KV cache between tenants.
Faster startup, no engine freezes
- Model weights load about six times faster. The image build used to replace the NCCL library of the vLLM base image with an older one, and loading GLM-5.3 took 266 seconds. The connector and SDK are now installed without touching the base image's packages, and the build fails if any of them moves: weights load in about 45 seconds.
- A lost gateway no longer freezes generation. Before, if the storage gateway went away, the engine could wait on it indefinitely while its health check still answered. The SDK client now gives up on a gateway that is gone and reconnects cleanly. With the gateway stopped under load in our test, generation never paused for more than 24 seconds and no request failed.
- Traceable builds. Every client kit now records the commit it was built from.
A full regression before release
The release passed a full regression on 8× NVIDIA H200 with the ADP accelerator, on GLM-5.3, GLM-5.2, MiniMax M3, DeepSeek-V4-Flash and Kimi-Linear, and on eleven configurations on A100: zero mismatches in the digests of restored KV pages (1.2 million pages on GLM-5.3 alone), no engine crashes, and a vLLM restart under load with long sessions running. A 196,000-token prompt on GLM-5.3 comes back from the KV cache in 2.12 seconds, slightly faster than the previous build.
One known issue is documented: hybrid models (Mamba or GDN layers) on vLLM 0.29 with
--mamba-cache-mode=align need --watermark=0.02 until a fix to vLLM's block accounting
lands.
Running a shared inference service?
Talk to our team about isolating tenants' KV caches and about the model, context lengths and concurrency your service needs.