Documentation menu
Overview

ADP KV cache documentation

The ADP KV cache connector keeps the KV cache of vLLM on the Awide Data Processor (ADP) card and brings it back on a prefix-cache hit, after eviction from GPU memory or a restart of the model server. This section covers how it works, what it runs on, and how to install and configure it.

What it does

A model server keeps the KV cache of every prompt it has seen in GPU memory, and only for as long as that memory lasts. When a long conversation, an agent’s context or a shared system prompt is evicted, or when the server restarts, the next request that starts the same way pays the full prefill again: tens of seconds for a long context on a large model.

The ADP KV cache connector gives vLLM a second, much larger tier for that cache. Every KV page the engine computes is also written to the ADP card; when a prompt comes back, vLLM restores its prefix from the card instead of recomputing it. Pages are keyed by vLLM’s own block hashes, so the cache survives restarts of the model server and is shared by every request with the same prefix. A 196,000-token prompt on GLM-5.3 comes back in 2.1 seconds instead of 47.7 seconds of prefill (release v6.1.1).

What you get:

  • Restore instead of recompute — the time to first token of a returning long prompt drops by an order of magnitude, and the GPUs spend that time on new work.
  • A cache that outlives the GPU — terabytes of KV cache on NVMe behind the card, kept across restarts, upgrades of the model server and evictions from GPU memory.
  • No extra host RAM tier — the only host memory the connector takes is a shared-memory transfer window of a few GiB; layers are restored one at a time while the model computes.
  • Plain vLLM — the connector plugs into vLLM’s KV connector API with one command-line option; no fork of the engine.

Architecture

The connector runs inside vLLM, next to every tensor-parallel rank. It hands KV pages to the storage gateway, a separate process on the same server, through a ring in shared memory. The gateway schedules the reads and writes, splits large objects into fragments the card accepts and spreads them over the card’s databases. The ADP card stores them on NVMe SSDs.

How the pieces fit
one GPU server · the ADP card local or remote
GPU SERVER vLLM engine one process per tensor-parallel rank GPU 0 GPU 1 GPU 2 … GPU N ADP KV cache connector · vLLM KV connector API KV pages, layer by layer Shared-memory ring /dev/shm on the host · a transfer window, not a cache read and written by slot Storage gateway schedules I/O, fragments large objects local · PCIe ADP card · in the GPU server PCIe Gen5 x8 slot, NVMe SSDs behind it Clients OpenAI-compatible API prompts OR: STORAGE SERVER ADP card · remote exported as an NVMe-oF target NVMe SSDs the capacity of the cache Either placement: the same connector and gateway, another gateway backend. NVMe-oF RDMA

The card can sit in either of two places, and nothing above the gateway changes between them:

PlacementWhere the card isHow the gateway reaches itGateway backend
Locala PCIe slot of the GPU serverthe ADP software stack on the same hostlocal-xdp
Remotea storage serverNVMe over Fabrics on an RDMA network (RoCE v2 or InfiniBand)remote-xdp

A remote card keeps the GPU servers free of storage and lets the cache sit on hardware sized for it; a local card needs no fabric. Architecture follows a page through both paths.

Compatibility at a glance

Supported
Inference enginevLLM 0.26, 0.28 and 0.29 — one connector code base; release images for 0.28.0 and 0.29.0
GPUsany NVIDIA GPU the vLLM image supports (compute capability 7.5 and newer in the upstream images)
CUDAthe CUDA of the vLLM base image: 13.0 for the default vllm/vllm-openai tags, 12.9 for the -cu129 tags
Model architecturesfull attention, sliding window, MLA, DeepSeek Sparse Attention, Mamba / Mamba-2 / Gated DeltaNet / Kimi Delta Attention hybrids, MoE; multi-token prediction (MTP)
Parallelismtensor parallelism; pipeline parallelism 1 for hybrid models
HostLinux x86_64, Docker or Kubernetes

The models verified on hardware, with the vLLM version each is recommended on, are listed in Supported models. Full requirements: Requirements.

Integration options

There are three ways to put the ADP card under vLLM. They share the gateway and the card; they differ in how the engine hands its KV cache over. Choosing an integration compares them.

Next steps

  1. Check the requirements for the GPU server, the card and the network.
  2. Set up the ADP card, local or remote.
  3. Install the gateway and vLLM with the connector, and verify a restore.
  4. Tune with the configuration reference and the notes for your model.
  5. Serving several tenants from one deployment? Set up tenant isolation.

Planning a deployment?

Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.

Talk to Awide Labs