On Kubernetes the gateway runs as a second container in the vLLM pod. Everything from Installation carries over; what changes is how the two containers share memory and how they are restarted. Any operator or chart that produces the pod below works the same way.
What the pod needs
| Requirement | How |
|---|---|
| The gateway next to vLLM | A second container in the same pod: one gateway per vLLM instance. |
One /dev/shm for both | hostIPC: true on the pod and no emptyDir (or any other volume) mounted at /dev/shm in either container. A private /dev/shm shows up in vLLM as Failed to open KVDatabase while the gateway reports the memory allocated. |
| A unique gateway port per node | The ring is a file in the node’s /dev/shm named after the gateway port, so two pods on one node need different --port values (and the matching xkv_gateway_port). |
| Access to the card | The gateway container runs privileged to reach the card’s device (/dev/pliops0 for a local card, the connected NVMe-oF namespace for a remote one) and send it pass-through commands. A remote card is connected on the node, outside the pod (ADP card setup). |
| The control port on loopback | Containers of one pod share the network namespace: bind the gateway to 127.0.0.1 (--bind-address=127.0.0.1). The SDK connects to localhost by default. |
| A fixed hash seed | PYTHONHASHSEED=0 in the vLLM container. |
| Pulling from the registry | An image pull secret for nexus.awide.io. |
| Scheduling | The pod must land on a node that has the card (local) or the connected namespace (remote): a node selector or affinity on a label you set on those nodes. |
Example
kubectl create secret docker-registry awide-registry \
--docker-server=nexus.awide.io --docker-username=<user> --docker-password=<token>
apiVersion: apps/v1
kind: Deployment
metadata:
name: glm53-adp
spec:
replicas: 1
selector:
matchLabels: { app: glm53-adp }
template:
metadata:
labels: { app: glm53-adp }
spec:
hostIPC: true # the ring lives in the node's /dev/shm
nodeSelector:
example.com/adp-card: "true" # a label you put on the nodes with the card
imagePullSecrets:
- name: awide-registry
containers:
- name: gateway
image: nexus.awide.io/awide-docker-hosted/akv-gateway:6.1.1
args:
- --storage-backend=remote-xdp
- --storage-devices=/dev/nvme1n1
- --fragment-size=131072
- --bind-address=127.0.0.1
- --port=5557
- --shared-memory-size=10GiB
- --thread-pool-size=64
- --num-dbs=128
- --start-clean=0
securityContext:
privileged: true # the card's device and NVMe pass-through
- name: vllm
image: nexus.awide.io/awide-docker-hosted/awideai-cuda-vllm-openai:v0.28.0-6.1.1
args:
- --model=/models/GLM-5.3-FP8
- --tensor-parallel-size=8
- --enable-prefix-caching
- --enable-prompt-tokens-details
- --kv-cache-dtype=fp8
- --max-num-batched-tokens=8192
- '--kv-transfer-config={"kv_connector":"FusIOnXConnector","kv_connector_module_path":"xkv_vllm_connector.fusionx_connector","kv_role":"kv_both","kv_connector_extra_config":{"xkv_gateway_port":5557,"chunk_size":128,"slot_alignment_bytes":4096}}'
env:
- { name: PYTHONHASHSEED, value: "0" }
ports:
- containerPort: 8000
readinessProbe:
httpGet: { path: /health, port: 8000 }
periodSeconds: 10
resources:
limits:
nvidia.com/gpu: 8
volumeMounts:
- { name: models, mountPath: /models, readOnly: true }
volumes:
- name: models
hostPath: { path: /models }
For a local card, the gateway container takes the local-card gateway image delivered with the ADP software stack,
--storage-backend=local-xdp instead of the three NVMe-oF options, and a hostPath volume for /var/log/pliops. The
gateway and vLLM images always come from the same release.
Check the shared memory once after the first start:
kubectl exec deploy/glm53-adp -c gateway -- sh -c 'echo ok > /dev/shm/adp-shm-check'
kubectl exec deploy/glm53-adp -c vllm -- cat /dev/shm/adp-shm-check # must print "ok"
kubectl exec deploy/glm53-adp -c gateway -- rm /dev/shm/adp-shm-check
Restarts
The gateway and vLLM form one unit: vLLM opens its connection to the gateway when it starts. Restart them together.
- When the gateway container restarts, recreate the pod (
kubectl delete pod …), so that vLLM connects to the new gateway. Until then vLLM serves from GPU memory alone and logsExternal KV cache turned off on this rank. - When the vLLM container restarts on its own, with the gateway running, nothing else is needed — that is exactly the restore path of Installation.
Recreating the pod on a restart of the gateway container is easy to automate: alert on the restart count of the
gateway container or on the log line above.
Security
The gateway’s control port answers lookups and deletes without authentication, and /metrics and the KV events of
vLLM show cache behaviour. Keep the gateway on loopback as above, and do not expose /metrics, KV events or the
gateway port outside the pod. For multi-tenant serving, see
Tenant isolation.
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.