NOTE
This backend is available for evaluation. It is not part of the product image or the client kit, and it is not yet recommended for high-concurrency serving with tensor parallelism (current limits). For production traffic use the vLLM KV connector.
What it is
vLLM has a model-agnostic KV offloading framework (OffloadingConnector, vllm/v1/kv_offload). In it, vLLM itself
handles everything that depends on the model — KV cache groups of hybrid models, page layouts, sliding windows, Mamba
state alignment, which caches are transferred at all — and hands a backend a key and the bytes of a chunk. The ADP
backend (XkvOffloadingSpec) stores those bytes on the card through the storage gateway.
TIP
The strength of this mode: any model, any vLLM version. Everything model-specific lives in vLLM, so the backend works with every model the offloading framework handles and on every vLLM release that ships it — a new architecture or a new vLLM version needs no work on our side. The backend detects the framework’s features rather than version numbers, and its contract is checked against the vLLM 0.26, 0.28 and 0.29 sources.
The price is that the framework loads the whole stored prefix before the step computes, instead of layer by layer under the computation, and that the framework has no way to report a failed transfer (current limits).
How it differs from the connector:
- One host copy. Pages are gathered on the GPU into a staging slot and copied straight into the gateway ring; a slot lives for one transfer, so there is no host-memory tier, index or eviction of its own.
- Tensor parallelism. For MLA models, whose page is the same on every rank, one object serves the whole group:
rank 0 writes, every rank reads, in lockstep through one gateway client group. Models with sharded or recurrent
groups above one GPU use one object per rank and a gateway port per rank (
tp_layout: per_rank). - Keys carry the model, the KV dtype, the rank layout and the byte layout of the page, so an object written under another configuration is a miss, never someone else’s bytes.
Enable it
On an image with vLLM 0.28.0 or 0.29.0, the backend package and its engine (delivered by Awide Labs for evaluation):
--kv-transfer-config '{"kv_connector":"OffloadingConnector","kv_role":"kv_both",
"kv_connector_extra_config":{
"spec_name":"XkvOffloadingSpec","spec_module_path":"xkv_offload",
"blocks_per_chunk":16,
"xkv.port":5557,"xkv.fragment_bytes":261120}}'
blocks_per_chunk sets how many KV blocks form one offloaded chunk. The gateway and the card are set up as for the
connector, and PYTHONHASHSEED=0 applies here too.
Settings
The backend’s own settings take the xkv. prefix in kv_connector_extra_config; every key, with its default, is in
Configuration → OffloadingSpec backend.
Current limits
- Concurrency with tensor parallelism. In replicated mode every rank must send the same requests in the same order. At high concurrency on many GPUs the ranks can drift out of that order, and a request then waits for a timeout. Single-GPU deployments are not affected.
- Restore speed. The framework loads the whole stored prefix before the tail of the prompt is computed. On large MLA models the faster read hides it; on hybrid models on one GPU a restore is noticeably slower than through the connector.
- Failed transfers. vLLM’s offloading framework has no way to report a failed transfer: a failed store is simply an object the card does not hold, and a failed load ends the request with an explicit error rather than serving stale pages.
- One data session per gateway. At most one engine uses a gateway at a time.
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.