Documentation menu
Get started

Installation

From a prepared ADP card to vLLM restoring prompts from it: get the images, run the storage gateway and vLLM with the connector side by side, and verify a restore.

This page installs the recommended integration, the vLLM KV connector, with Docker Compose on one GPU server. Kubernetes takes the same pieces; the LMCache plugin and the OffloadingSpec backend have pages of their own.

1. Prepare the card

Set up the card as described in ADP card setup and check that pliocli system get_status reports Aggregated status is OK and key eviction is enabled. Then note the device the gateway will use:

PlacementDevice for the gateway
Local card/dev/pliops0 (card 0), and the log directory /var/log/pliops
Remote cardthe NVMe-oF namespace connected on the GPU server, for example /dev/nvme1n1 (nvme list, model XKV_LIGHTNING_AI)

2. Get the software

The images and client kits are published in the Awide Labs registry. The account is issued on request — contact us to get one.

docker login nexus.awide.io
ArtifactCoordinates
vLLM 0.29.0 with the connectornexus.awide.io/awide-docker-hosted/awideai-cuda-vllm-openai:v0.29.0-6.1.1
vLLM 0.28.0 with the connectornexus.awide.io/awide-docker-hosted/awideai-cuda-vllm-openai:v0.28.0-6.1.1
Storage gateway (remote card)nexus.awide.io/awide-docker-hosted/akv-gateway:6.1.1
Client kit for vLLM 0.29.0https://nexus.awide.io/repository/awide-ai-bundles/awideai/akv-kit-cuda-v0.29.0-6.1.1.tar.gz
Client kit for vLLM 0.28.0https://nexus.awide.io/repository/awide-ai-bundles/awideai/akv-kit-cuda-v0.28.0-6.1.1.tar.gz

Pick the vLLM version your model is recommended on (Supported models). Use the gateway and the vLLM image (or kit) of the same release: at the handshake the gateway refuses an SDK of another version (VERSION_MISMATCH in its log), and the connector then stays off. The gateway with the local-card backend is delivered together with the ADP software stack. The product image is the upstream vllm/vllm-openai image of that version with the SDK and the connector added; nothing else in it is changed, and it starts the OpenAI-compatible server like the upstream image does.

docker pull nexus.awide.io/awide-docker-hosted/awideai-cuda-vllm-openai:v0.28.0-6.1.1
docker pull nexus.awide.io/awide-docker-hosted/akv-gateway:6.1.1

Building on your own vLLM image. If you run a vLLM image of your own — a CUDA 12.9 variant, a build with your patches — use the client kit instead of the product image. It compiles the SDK against the Python, PyTorch and CUDA of that image and installs the connector on top, without changing the image’s packages:

curl -fu <user> -O https://nexus.awide.io/repository/awide-ai-bundles/awideai/akv-kit-cuda-v0.28.0-6.1.1.tar.gz
tar xzf akv-kit-cuda-v0.28.0-6.1.1.tar.gz && cd akv-kit-cuda-v0.28.0-6.1.1
./build.sh --base-image vllm/vllm-openai:v0.28.0-cu129 --tag my-registry/vllm-openai-awideai:6.1.1

3. Write the compose file

The gateway and vLLM run as two containers. The gateway creates the shared-memory ring as a file in /dev/shm, so both must use the host’s /dev/shm: ipc: host on both, and nothing mounted over /dev/shm in either. vLLM reaches the gateway’s control port by its service name (XKV_GATEWAY_HOSTNAME).

The file below is for a remote card and GLM-5.3 FP8 on eight GPUs; the values that depend on the model are marked, and Supported models has them for the other models.

services:
  gateway:
    image: nexus.awide.io/awide-docker-hosted/akv-gateway:6.1.1
    ipc: host                          # the ring is a file in the host's /dev/shm
    cap_add: [SYS_ADMIN, IPC_LOCK]     # SYS_ADMIN: NVMe pass-through commands to the card
    ulimits:
      memlock: -1
    devices:
      - /dev/nvme1n1:/dev/nvme1n1      # the connected ADP namespace (nvme list)
    command:
      - --storage-backend=remote-xdp
      - --storage-devices=/dev/nvme1n1
      - --fragment-size=131072         # the NVMe-oF transfer size of the card
      - --port=5557
      - --shared-memory-size=10GiB
      - --thread-pool-size=64
      - --num-dbs=128
      - --start-clean=0                # keep the cache across restarts
    restart: unless-stopped

  vllm:
    image: nexus.awide.io/awide-docker-hosted/awideai-cuda-vllm-openai:v0.28.0-6.1.1
    ipc: host
    cap_add: [IPC_LOCK]
    ulimits:
      memlock: -1
    depends_on: [gateway]
    ports:
      - "8000:8000"
    deploy:
      resources:
        reservations:
          devices:
            - { driver: nvidia, count: all, capabilities: [gpu] }
    environment:
      PYTHONHASHSEED: "0"              # required: otherwise nothing is found after a restart
      XKV_GATEWAY_HOSTNAME: gateway    # where the SDK finds the gateway
      # VLLM_USE_V2_MODEL_RUNNER: "0"  # required on vLLM 0.29
    volumes:
      - /models:/models:ro
    command:
      - --model=/models/GLM-5.3-FP8
      - --tensor-parallel-size=8
      - --enable-prefix-caching
      - --enable-prompt-tokens-details
      - --kv-cache-dtype=fp8                 # model-specific
      - --max-num-batched-tokens=8192        # model-specific
      - '--kv-transfer-config={"kv_connector":"FusIOnXConnector","kv_connector_module_path":"xkv_vllm_connector.fusionx_connector","kv_role":"kv_both","kv_connector_extra_config":{"xkv_gateway_port":5557,"chunk_size":128,"slot_alignment_bytes":4096}}'
    restart: unless-stopped

Keep the JSON of --kv-transfer-config (and of --speculative-config, if you use MTP) in single quotes as above, and add the model’s own options — tool-call and reasoning parsers, --max-model-len — as usual.

For a local card, only the gateway changes: the gateway image with the local backend, the card’s device and log directory instead of the NVMe namespace, and the local backend’s default fragment size.

  gateway:
    image: <local-card gateway image, release 6.1.1>   # delivered with the ADP software stack
    ipc: host
    cap_add: [IPC_LOCK]
    ulimits:
      memlock: -1
    devices:
      - /dev/pliops0:/dev/pliops0
    volumes:
      - /var/log/pliops:/var/log/pliops
    command:
      - --storage-backend=local-xdp    # fragment size defaults to 261120 here
      - --port=5557
      - --shared-memory-size=10GiB
      - --thread-pool-size=64
      - --num-dbs=128
      - --start-clean=0
    restart: unless-stopped

On a local card the connector needs no slot_alignment_bytes. Every gateway setting can also come from a YAML file (--config /etc/akv/gateway.yaml); flags given on the command line win. All of them: Configuration.

4. Start

docker compose up -d
docker compose logs -f gateway vllm
curl -f http://127.0.0.1:8000/health

The first start of a large model takes a while: weights load, CUDA graphs are captured. When the connector has registered, the vLLM log shows which caches it exchanges and how:

XKV: <N> cache tensors registered for exchange
KV cache registry: <N> cache(s) in <G> group(s); writers: layer_hook=... ; readers: ...

A model whose caches the connector cannot account for does not start at all, and the log names the cache — the connector never runs half-configured.

5. Verify a restore

A restore is proved by a hit that can only come from the card: send a long prompt, restart only vLLM (the gateway and the card keep the cache), and send it again.

MODEL=/models/GLM-5.3-FP8
PROMPT=$(python3 -c "print(' '.join(f'Line {i}: the ADP card keeps this prompt.' for i in range(3000)))")
ask() {
  curl -s http://127.0.0.1:8000/v1/completions -H 'Content-Type: application/json' \
    -d "$(jq -n --arg m "$MODEL" --arg p "$PROMPT" '{model: $m, prompt: $p, max_tokens: 8}')" \
    | jq '.usage | {prompt_tokens, cached: .prompt_tokens_details.cached_tokens}'
}

ask                                # cold: cached is 0, the prompt is computed and written to the card
docker compose restart vllm        # GPU memory is empty now, the card is not
until curl -sf http://127.0.0.1:8000/health; do sleep 10; done
time ask                           # warm: cached is close to prompt_tokens, and the answer comes much faster

cached_tokens is rounded down to whole chunks: with chunk_size 128 a prompt of 30,100 tokens can hit at most 30,080. A prompt shorter than one chunk is never stored. Over time the vLLM log reports the hit rate of the card, and /metrics exposes the same counters:

External prefix cache hit rate: 99.6%
vllm:external_prefix_cache_queries_total   vllm:external_prefix_cache_hits_total

If cached_tokens stays at 0 after the restart, see Operations & troubleshooting.

Stopping and upgrading

docker compose down keeps the cache on the card: with --start-clean=0 the next start finds it. An upgrade of the images keeps it too, unless the release notes say the stored layout changed — then wipe the card as described in ADP card setup. Changing the model, chunk_size or the tensor-parallel size starts a new cache in any case (Configuration).

Planning a deployment?

Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.

Talk to Awide Labs