What runs on the card’s host
The card is managed by the ADP software stack on the server it sits in — the GPU server for a local card, a storage
server for a remote one: a driver, one storage service per card (pliostore@<N>), the pliocli tool and, for a remote
card, one NVMe-oF target service per card (lightning-spdk-target-<N>). What each part is and what the host needs for
it: ADP card requirements.
Install the software stack
The ADP software stack is delivered by Awide Labs as an installer built for the kernel of the card’s host. What the host needs before you start — slot, SSDs, packages, kernel — is in ADP card requirements.
-
Install the card in a PCIe Gen5 x8 slot and the NVMe SSDs that will hold the cache in the same server, on the card’s NUMA node.
-
Install the prerequisite packages.
-
Run the installer with the SSDs of the card listed explicitly. It installs the driver, the storage service and the command-line tool, puts the SSDs into one RAID 0 array behind the card and reserves hugepages:
sudo ./xdp-installer install --level=0 --resources="/dev/nvme1n1 /dev/nvme2n1" --vhc=0 \ --enable-set-hugepages --hugepage-distribution-percentage=50 --assume-yes -
Remote card only: install the NVMe-oF target on the address the GPU server will connect to, with the 128 KiB I/O size the gateway’s fragments are sized for, and start it:
./lightning_ai_spdk_target.py install --targets=192.0.2.10:4420 --start-service --io-size 131072 -
On a storage server, set the CPU and I/O tuning profile:
sudo tuned-adm profile throughput-performance -
Check that the card is up (next section), then configure the storage service.
Check the card
systemctl is-active pliostore@0 # the storage service of card 0: "active"
pliocli system get_xdp_list # the cards on this host, their device paths and mode
pliocli system get_status # every component, and the aggregated status
pliocli system get_disk_usage -x 0 # physical usage, valid objects, databases
pliocli raid get_raid_info -x 0 # the SSD array behind card 0
systemctl is-active lightning-spdk-target-0 # remote card only: the NVMe-oF target of card 0
A healthy card reports every component OK:
Components status:
Driver (OK): Driver is working properly
Service (OK): Pliostore service is working properly and system is active
RAID (OK): RAID Array is functioning properly, raidLevel: RAID0, protection: No, performance: Normal
CapacityUsage (OK): Capacity usage is below defined thresholds
Firmware (OK): Firmware is working properly
Temperature (OK): Temperature is at NORMAL level
SuperCapacitors (OK): Super Capacitors are working properly
OnboardFlash (OK): Onboard flash component is working properly
Aggregated status is OK
The usage report shows how full the card is and how many objects and databases it holds; the number of databases is
the gateway’s num_dbs (128 by default) once a gateway has connected:
Physical usage [%] : 12.40
Physical capacity : 863.10 GB / 6.96 TB
Number of valid objects : 10051238 / 1680330742
Number of databases : 128 / 4096 (0 deleting)
Configure the storage service
The storage service of card N reads /etc/pliops/<N>/pliostore.ini, which the installer writes.
Check two keys on every card that serves the KV cache — the full list is in
Configuration → ADP storage service:
[db_settings]
key_eviction = true
product_type = 0 # 0 = KV, 1 = Block
Apply a change with the model server and the gateway stopped:
systemctl restart pliostore@0
journalctl -u pliostore@0 -o cat | grep -i "key eviction" # "Key eviction Enabled"
systemctl restart lightning-spdk-target-0 # remote card: restart the target after the service
A remote card then has to be reconnected on the GPU server.
Capacity and key eviction
The storage service reports usage in three zones, visible as CapacityUsage in pliocli system get_status:
| Physical usage | Status | Effect |
|---|---|---|
| below the soft threshold (90 % by default) | OK | — |
| between the soft and the hard threshold (100 %) | warning | none on writes |
| above the maximum | CRITICAL | every incoming write fails until space is freed; the service logs event 3002 |
With key_eviction = true the card starts reclaiming space before it reaches the maximum. The point where eviction
starts is built into the storage service, not a setting in the INI file, and the soft and hard thresholds
(pliocli system get_usage_soft_threshold, get_usage_hard_threshold) only change the reported status — do not move
them expecting eviction to start earlier.
Watch the physical usage (pliocli system get_disk_usage) like the fill level of any cache, and alert well before the
maximum: a card that fills faster than it reclaims shows up as a falling hit rate, not as an error in vLLM.
Wiping the card
A wipe empties the cache: after a change of the stored layout (settings that require a clean
cache) it is required, otherwise it only costs a refill. The
gateway does it at start with --start-clean=1, by deleting its databases — which the storage service refuses while key
eviction is on (Key eviction is ENABLED, delete DB not supported). So a wipe takes this order:
- Stop vLLM and the gateway.
- Set
key_eviction = falseand restart the storage service (and the target, then reconnect — remote card). - Start the gateway once with
--start-clean=1and wait until its log reports the databases deleted;pliocli system get_disk_usage -x 0then shows0objects. - Set the gateway back to
--start-clean=0. - Set
key_eviction = trueand restart the storage service (and the target, then reconnect). - Start the gateway and vLLM.
WARNING
With key eviction enabled, keep start_clean at 0 everywhere the gateway is started — compose files, Kubernetes
manifests, restart scripts. A gateway that starts with start_clean=1 on every restart would otherwise try to wipe
the card each time.
Remote card: connect the GPU server
The NVMe-oF target service on the storage server exports one key-value namespace per card over RDMA, port 4420. On
the GPU server, connect it with the kernel initiator:
modprobe nvme_rdma
nvme discover -t rdma -a 192.0.2.10 -s 4420 # lists the subsystem NQN of the card
nvme connect -t rdma -a 192.0.2.10 -s 4420 -n <subsystem NQN>
nvme list # the ADP namespace, model XKV_LIGHTNING_AI
nvme list-subsys # ... rdma traddr=192.0.2.10,trsvcid=4420 live
The namespace is a key-value namespace, not a block device: nvme list shows its size as 0 B and its SMART log is
empty — both are normal. Do not partition or format it. Its device path (/dev/nvme1n1, for example) is what the
gateway gets as --storage-devices, and the gateway container needs that device and SYS_ADMIN for the NVMe
pass-through commands (Installation).
Make the connection persistent the way your distribution does it (for example /etc/nvme/discovery.conf and the
nvmf-autoconnect service of nvme-cli), and check the device name after every reconnect: it can change.
After a restart of the target
A restart of the storage service or of the target on the storage server can change the namespace identifiers, and the connected device stops answering. On the GPU server:
nvme disconnect -n <subsystem NQN>
nvme connect -t rdma -a 192.0.2.10 -s 4420 -n <subsystem NQN>
nvme list # the device path may differ from before
Then restart the gateway and vLLM together, so that both open the device again.
Transfer size and fragments
| Placement | Limit | Gateway fragment_size |
|---|---|---|
| Local card | an object of at most 256 KiB | 261120 (255 KiB; the default of the local backend) |
| Remote card | one NVMe command of at most MDTS — 128 KiB with the target’s --io-size 131072 — and a multiple of 4 KiB | 131072 |
The gateway logs the device limit at start (# Max I/O Size) and refuses a fragment size that does not fit it.
Models whose KV page is not a multiple of 4 KiB (GLM-5.x) also need slot_alignment_bytes: 4096 in the connector
settings on a remote card. Changing fragment_size requires a wipe.
Durability
The card holds a cache, and the system is built to lose nothing that matters if it loses some of it: a write the card has acknowledged can still be missing after an abrupt restart of the storage service or the target, and the connector then reads that chunk as a miss and vLLM recomputes it. The card itself protects its data against power loss with on-board supercapacitors.
Reading the gateway log
NVMe media error type: 2 code: 135 is how the card answers a lookup of a key it does not hold — an ordinary cache
miss, which the gateway logs at ERROR level. Thousands of them during a run are normal. The same code is also returned for other
refusals, a full card among them, so judge the card by pliocli system get_status and get_disk_usage, not by the
count of these lines.
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.