Hardware
The card
| Requirement | |
|---|---|
| Slot | PCIe Gen5, 8 lanes (the card is a PCIe Gen5 x8 device). |
| Form factor | Low-profile, half-height half-length (HHHL, 6.6” × 2.536”); with tall and short brackets for full-height and low-profile slots. |
| Power | Under 25 W typical, 45 W at most, +12 V through the PCIe slot — no auxiliary power cable. |
| Cooling | Operating temperature 10–52 °C at 250 LFM of airflow across the card; storage 5–35 °C, below 90 % humidity, non-condensing. |
| Power-loss protection | On-board supercapacitors; no battery and nothing to install. |
| Servers | Any standard x86_64 server with a free slot: Dell, HPE, Lenovo, Supermicro, Quanta, Wiwynn, Inspur, Sugon, Fujitsu, Hitachi, Tyan, MiTAC, Intel, Cisco, AIC. |
GPU servers often run their fans by GPU temperature; check that the slot you choose gets the airflow above when the GPUs are idle too.
The SSDs behind the card
The card stores the cache on NVMe SSDs in the same server. At installation the SSDs are assigned to the card and combined into one array behind it.
| Requirement | |
|---|---|
| Interface | NVMe (PCIe Gen3, Gen4 or Gen5). The card also drives SAS and SATA SSDs; for the KV cache use NVMe. |
| Type and vendor | TLC or QLC flash from any vendor (Samsung, WD, Micron, Intel, Kioxia, Hynix, Seagate, …). |
| Array | RAID 0 across the SSDs of one card. The content is a cache that can always be recomputed, so there is no need to pay for redundancy. |
| Capacity | Up to 128 TB of raw SSD capacity per card. The usable cache is about 90 % of the raw capacity: two 3.84 TB SSDs give about 7 TB. |
| Dedicated | The SSDs belong to the card: no file system, no partitions, not the boot disk. Keep the operating system on a disk of its own. |
| Placement | On the same NUMA node (CPU socket) as the card. |
How much capacity. Plan from the prefixes you want to keep warm, not from the request rate: the KV size per token
of your model times the tokens of all long prompts, conversations and agent contexts that should come back without a
prefill. The usage report of the card (pliocli system get_disk_usage) shows the stored objects and their average size
once a model has run for a while. When the card is full it evicts on its own and the hit rate is what the capacity
allows (capacity and key eviction).
Where the card goes
| Placement | Hardware |
|---|---|
| Local — in the GPU server | A free PCIe Gen5 x8 slot in the GPU server and room for the SSDs there. No network. |
| Remote — in a storage server | A storage server with the card and its SSDs (below), and an RDMA network between it and the GPU server. The GPU server needs only the network card. |
A remote card keeps the GPU servers free of storage and lets the cache sit on hardware sized for it; a local card is the simplest setup, with nothing between the gateway and the card.
Several cards in one server
Put one card per NUMA node: in a two-socket server, a card on each socket, each with its own SSDs on that socket. Every card has its own storage service and, in a remote setup, its own NVMe-oF target and network address. Give each gateway a card of its own; several vLLM instances on one GPU server can each use a separate card.
Storage server for a remote card
| Requirement | |
|---|---|
| CPU | x86_64. The NVMe-oF target service polls its CPU cores continuously, so give every card a set of dedicated cores on its own NUMA node — they show as 100 % busy, which is normal. |
| Memory | Enough for the operating system plus the hugepages the installer reserves for the target service. |
| Network | An RDMA-capable NIC per card, on the card’s NUMA node (below). |
| Disks | The SSDs of the cards, and a separate boot disk. |
| Remote management | IPMI or another out-of-band console is strongly recommended: the target runs on bare metal. |
Network for a remote card
| Requirement | |
|---|---|
| Transport | NVMe over Fabrics on RDMA: RoCE v2 over Ethernet, or InfiniBand. |
| NIC | RDMA-capable, on the storage server and on the GPU server; on the storage server on the same NUMA node as its card. |
| Speed | 100 Gb/s or faster per card. The restore speed of a long prompt is capped by the slower of the link and the SSDs. |
| MTU | Jumbo frames, 9000, end to end — NICs and switch ports alike. |
| Addressing | One IP address and port (4420) per card’s target; the GPU server must reach it, tenants must not. |
A slower link works, just with slower restores; for RoCE, configure the switches the way your NIC vendor recommends for RDMA traffic.
Software
The ADP software stack
Everything that runs the card comes in one package from Awide Labs, the ADP software stack. It is installed on the host of the card — the GPU server for a local card, the storage server for a remote one.
| Component | What it is | On the host as |
|---|---|---|
| Driver | The kernel module of the card. Built for a specific kernel version. | kernel module |
| Firmware | The card’s firmware, delivered with the stack. | reported by pliocli system get_status |
| Storage service | Keeps the key-value databases on the card; one instance per card. Starts at boot. | pliostore@<N> (systemd), config /etc/pliops/<N>/pliostore.ini |
| Command-line tool | Status, usage, the SSD array, maintenance. | pliocli |
| NVMe-oF target service (remote card only) | Exports the card’s key-value namespace over RDMA, one instance per card. | lightning-spdk-target-<N> (systemd), port 4420 |
| Installers | Install the above and build the SSD array; the target installer sets up the NVMe-oF target. | xdp-installer, lightning_ai_spdk_target.py |
On the GPU server, the connector and the gateway come as container images (Installation). With a local card the gateway container talks to the storage service on the same host; with a remote card nothing of the stack runs on the GPU server — only the kernel’s NVMe-oF initiator.
Release 6.1.1 of the connector is verified with ADP software stack 2.4.5.
Operating system and kernel
| Host | Requirement |
|---|---|
| Host of the card | Linux x86_64. The driver is built for the exact kernel the host runs: tell Awide Labs your distribution and kernel version and you get the stack built for it. A kernel update needs a stack built for the new kernel — pin the kernel on these hosts. |
| GPU server, remote card | Linux x86_64 with the NVMe-oF RDMA initiator in the kernel (nvme_rdma) — any current distribution kernel has it. |
Packages to install first
On the host of the card the installers need these packages (names as in RHEL-family distributions; the Debian family has equivalents):
| Package | Version | Why |
|---|---|---|
nvme-cli, pciutils, libpciaccess | any | Finding and managing the card and its SSDs. |
libaio, numactl, numactl-devel | any | I/O and NUMA placement of the storage service. |
ledmon | 0.96 or newer | Drive LEDs of the SSD array. |
fuse3, fuse3-devel, libuuid-devel, ncurses-devel | any | Required by the stack’s services and tools. |
python3, python3-pip | 3.8 or newer | The installers; pip modules meson and pyelftools for the target installer. |
libibverbs, libibverbs-devel, librdmacm-devel | any | RDMA for the NVMe-oF target (remote card). |
tuned | any | The throughput-performance profile of a storage server. |
The target installer also installs the dependencies of the NVMe-oF target with RDMA support itself.
On the GPU server with a remote card:
| Package or setting | Why |
|---|---|
nvme-cli | nvme discover, nvme connect, nvme list. |
kernel modules nvme_rdma, rdma_ucm | The NVMe-oF initiator over RDMA. |
nvme_core.multipath=N (kernel parameter) | The gateway sends NVMe pass-through commands to the namespace device; native NVMe multipath must be off. |
the RDMA user-space stack of your NIC (rdma-core or the vendor’s driver package) | RDMA on the NIC. |
Installing and checking it
ADP card setup installs the stack, checks the card and its services, and sets the two storage service keys the KV cache depends on — key eviction on, key-value mode. Every key is listed in Configuration → ADP storage service.
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.