Hardware
| Requirement |
|---|
| Server | x86_64 with NVIDIA GPUs that vLLM supports. |
| GPUs | The upstream vLLM images are built for compute capability 7.5 and newer: Turing, Ampere, Ada, Hopper, Blackwell. Some models need more than that from vLLM itself — DeepSeek Sparse Attention models run only on Hopper and Blackwell. |
| GPU memory | What the model and its KV cache need without the connector, plus a small headroom: the connector’s buffers are allocated after vLLM sizes its KV cache (resources). |
| ADP card | Local placement: a free PCIe Gen5 x8 slot and the SSDs behind it, see ADP card hardware. Remote placement: nothing of the card here, only the network card below. |
| Network (remote card) | An RDMA-capable NIC (RoCE v2 or InfiniBand), 100 Gb/s or faster, jumbo frames (MTU 9000), see network. |
Operating system and software
| Requirement |
|---|
| OS | Linux x86_64 |
| NVIDIA driver | One that supports the CUDA of the vLLM image: R580 or newer for the CUDA 13.0 images (the default vllm/vllm-openai tags), R575 or newer for the CUDA 12.9 images (-cu129 tags). |
| Containers | Docker Engine 25 or newer with Compose v2 and the NVIDIA Container Toolkit (nvidia runtime), or Kubernetes with the NVIDIA device plugin. |
| Local card | The ADP software stack on this server (ADP card software). |
| Remote card | nvme-cli; the kernel modules nvme_rdma and rdma_ucm; the RDMA user-space stack of your NIC; NVMe native multipath off (kernel parameter nvme_core.multipath=N). |
Everything else — CUDA, PyTorch, Python, vLLM, the connector and the SDK — comes inside the container image. The
connector is installed into the vLLM image without changing any of its packages.
Resources
| Resource | What to plan for |
|---|
/dev/shm | At least the ring of every gateway on the host (shared_memory_size, 6 GiB by default, 10 GiB for 8-GPU servers), plus what vLLM uses itself. The gateway and vLLM must share the host’s /dev/shm. |
| Host RAM | The ring above plus vLLM’s own needs. The connector keeps no cache in host memory. |
| CPU | A few cores for the gateway: its worker threads (thread_pool_size, 4 by default, 64 on 8-GPU servers) each run one device command at a time. |
| Disk | About 50 GB for the images, plus the model weights. |
| GPU memory headroom | NVLink fan-out, when enabled, takes a 256 MiB staging buffer and an NCCL communicator per GPU outside gpu_memory_utilization; lower --gpu-memory-utilization by about 0.005 on a 141 GB GPU for the defaults. Long prompts on large models may need a little more headroom than without the connector. |
Software from Awide Labs
| Artifact | What it is |
|---|
| Product image | vLLM with the SDK and the connector installed, for vLLM 0.28.0 and 0.29.0 |
| Gateway image | The storage gateway for a remote card (and a directory backend for trials). The gateway with the local-card backend comes with the ADP software stack. |
| Client kit | SDK and connector sources with a Dockerfile, to build on your own vLLM image |
The images and kits are published in the Awide Labs registry nexus.awide.io; the account is issued on request. The
gateway and the vLLM image (or kit) must come from the same release: the gateway refuses an SDK of another version.
Installation has the exact names.
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.
Talk to Awide Labs