Documentation menu
Reference

Configuration

Every setting in one place: the storage gateway, the vLLM KV connector, environment variables, the vLLM options they depend on, the ADP storage service, the LMCache plugin, the OffloadingSpec backend, and the settings that require a clean cache.

Where settings live

ComponentSet inReference
Storage gatewaycommand-line flags, or a YAML file given with --configStorage gateway
vLLM KV connectorkv_connector_extra_config of --kv-transfer-config; per request kv_transfer_paramsConnector
vLLM processenvironment variablesEnvironment variables
vLLMits command linevLLM options
ADP storage service/etc/pliops/<N>/pliostore.ini on the card’s hostADP storage service
LMCache pluginadapter_params of the L2 adapterLMCache plugin
OffloadingSpec backendxkv.* keys in kv_connector_extra_configOffloadingSpec backend

Complete deployments — compose files and a Kubernetes pod — are in Installation and Kubernetes.

Storage gateway

The gateway takes its settings from the command line, from a YAML file (--config <file>), or both; a flag given on the command line wins over the file. The flag is the key with dashes: shared_memory_size → --shared-memory-size. Sizes take B, KiB, MiB, GiB, TiB (1024-based) or KB, MB, GB, TB (1000-based); a bare number is bytes.

# gateway.yaml — a remote card on an 8-GPU server
storage_backend: remote-xdp
storage_devices:
  - /dev/nvme1n1
fragment_size: 131072
port: 5557
shared_memory_size: 10GiB
thread_pool_size: 64
num_dbs: 128
start_clean: 0
log_severity: info
KeyDefaultWhat it does
storage_backendremote-xdpWhere objects go: remote-xdp (an ADP card over NVMe-oF), local-xdp (an ADP card in this server), file-system (a directory; for trying the software without a card), loopback (stores nothing; for tests).
storage_devices—Required for remote-xdp: the connected NVMe-oF namespace(s) of the card, for example /dev/nvme1n1. The gateway refuses to start without it.
fragment_size0 (off); 261120 on local-xdpObjects larger than this are stored as several fragments. Local card: 261120. Remote card: a multiple of 4 KiB no larger than the device’s transfer limit — 131072. At least 4096 when set. ⚠️ Changing it requires a clean cache.
fragment_spread_across_dbstruePut the fragments of one object into different databases, so they are read in parallel. ⚠️ Changing it requires a clean cache.
port5557The control port (ZeroMQ). Also names the ring file in /dev/shm, so gateways on one host need different ports.
bind_address*The address the port listens on: an IPv4 address, an interface name or *. The port answers lookups and deletes without authentication — bind it to 127.0.0.1 when vLLM shares the gateway’s network namespace (one pod, or host networking), and keep it off tenant networks otherwise.
shared_memory_size6GiB (at least 512MiB)The ring in /dev/shm, split between the tensor-parallel ranks of the engine. A transfer window, not a cache: its capacity is checked per request, so it does not limit the context length. 10 GiB is a good value for 8-GPU servers.
thread_pool_size4Worker threads; each runs one device command at a time, so this is also the most commands in flight. The main lever of restore speed: 64 on 8-GPU servers.
core_listall coresCPU cores to pin the workers to, for example "0-7,16-23".
num_dbs128Number of databases on the card the objects are spread over.
start_clean01 deletes the gateway’s databases at start, i.e. wipes the cache. Use once, deliberately; keep 0 while key eviction is enabled on the card (wiping the card).
scheduler_max_parallel_reqs2Read requests running at once. Keep 2, the value validated on hardware (the gateway warns above it); long restores get faster with thread_pool_size, not with this.
scheduler_overlap_trigger_ratio0.8 (1.0 on local-xdp)Share of the previous read that must be done before the next one may start.
sequence_timeout_sec30 (at most 33)How long a request sent by some ranks of an engine may wait for the others before it fails.
log_diremptyA directory for log files; empty logs to standard output only.
log_severityinfoperf, trace, debug, info, notice, warning, error or fatal.

Connector

The keys of kv_connector_extra_config in --kv-transfer-config. A boolean key also takes 1/0, "true"/"false", "yes"/"no", "on"/"off"; anything else stops the connector at start.

KeyDefaultWhat it does
xkv_gateway_port5557The gateway’s control port.
chunk_sizevLLM’s scheduler block sizeTokens per stored object; must be a multiple of the block size of every KV cache group (for sliding-window models the window must be at least twice chunk_size and a multiple of it). ⚠️ Changing it starts a new cache.
slot_alignment_bytes0Round each transfer up to this many bytes. 4096 on a remote card for models whose page is not a multiple of 4 KiB (GLM-5.x). ⚠️ Changing it requires a clean cache.
enable_deferred_writetrueBatch the writes of decode steps that run under full CUDA graphs into one request per step. false makes the connector ask vLLM for piecewise graphs (slower decode).
layer_callbacksautoWhether the engine’s layer hooks are used for this model; auto decides from the model and checks it against the registered layers.
aggregated_get_submission, aggregated_put_completiontrueSubmit reads and complete writes in batches.
synchronous_layer_getfalseMake every layer wait for its own read before computing even where the read path does not need it. A debugging guard.
verify_all_layersfalseConfirm a cached prefix on every exchanged layer instead of one layer per group.
clean_cachefalseWipe the cache when the connector starts. For tests.
log_levelINFORanks other than 0 log from NOTICE up unless a level is set.
num_copy_thread_block_get, num_copy_threads_per_block_get, num_copy_thread_block_put, num_copy_threads_per_block_putSDK defaultsCopy kernel launch parameters.
nvlink_fanoutfalseFan restores of MLA models out over NVLink (details).
nvlink_fanout_groupsautoNVLink groups of ranks, auto from NVML or explicit.
nvlink_fanout_staging_mb256Staging buffer per rank, MiB.
nvlink_fanout_min_layer_bytes4194304Smaller layers are copied whole by every rank.
key_namespaceemptyMix a deployment name (up to 64 bytes) into every key. ⚠️ Setting, changing or clearing it starts a new cache.
require_cache_saltfalseRequests without a cache_salt neither read nor store the external cache.
cache_rules—Experimental override of who writes a cache, by layer class. Not for production.

Per request, in the request body: "kv_transfer_params": {"xkv_store": false} keeps the request out of the external cache — no lookup, no restore, no write.

Environment variables

VariableWhereWhat it does
PYTHONHASHSEED=0vLLMRequired. A fixed hash seed, so block hashes — and the keys on the card — are the same after a restart.
XKV_GATEWAY_HOSTNAMEvLLMHost of the gateway’s control port; localhost by default. In Docker Compose, the gateway’s service name.
VLLM_USE_V2_MODEL_RUNNER=0vLLM 0.29Required on 0.29: the connector works with the V1 model runner.
NCCL_LAUNCH_ORDER_IMPLICIT=1vLLMRecommended with NVLink fan-out (NCCL 2.26 or newer); harmless otherwise.
XKV_EXIST_REOPEN_COOLDOWN_SvLLMSeconds between attempts to reopen the lookup connection after a timeout; 30 by default.
XKV_STEP_METRICS_EVERYvLLMReport per-step timing of the connector every this many steps; 500 by default.
XKV_SYNCHRONOUS_LAYER_GETvLLMOverrides synchronous_layer_get.
XKV_NVTXvLLMNVTX ranges around the connector’s work, for profiling.

vLLM options the connector depends on

OptionWhenWhy
--enable-prefix-cachingalwaysThe connector serves prefix-cache hits; hybrid models refuse to start without it.
--enable-prompt-tokens-detailsrecommendedReports cached_tokens in API responses.
--mamba-cache-mode alignMamba / GDN / KDA hybridsThe only Mamba cache mode the connector accepts; vLLM 0.29 picks it by itself with prefix caching.
--watermark=0.02Mamba / GDN / KDA hybridsKeeps the few free blocks a hybrid needs when a request’s prefix is restored, on a nearly full KV cache (details).
--kv-cache-dtype fp8DeepSeek Sparse Attention models, MiniMax M3One page size per group.
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}'DeepSeek Sparse Attention models on vLLM 0.29Under piecewise graphs the engine does not write their KV.
--max-num-batched-tokens=8192GLM-5.xAvoids an out-of-memory error in the sparse MLA kernel on prompts from about 128k tokens.
--no-async-schedulingfull + sliding-window models (Gemma 4, GLM-4.7-Flash) with deferred writeDeferred write of a sliding-window group needs synchronous scheduling and pipeline parallelism 1; the alternative is enable_deferred_write: false. The connector says so at start.
--gpu-memory-utilizationNVLink fan-outAbout 0.005 lower on a 141 GB GPU for the fan-out buffers.

ADP storage service

/etc/pliops/<N>/pliostore.ini on the host of card N, written by the installer. Apply a change with the gateway and vLLM stopped: systemctl restart pliostore@<N>, and on a storage server the target after it (ADP card setup).

SectionKeyValueWhy
[db_settings]key_evictiontrueRequired. Lets the card reclaim space on its own. Without it a full card refuses every write: vLLM keeps serving, but nothing new reaches the cache and the hit rate decays to zero. It is off unless set.
[db_settings]product_type0Key-value mode, which the gateway uses (1 is block mode).
[qos]aio_read_io_size, aio_write_io_sizeas delivered (131072)The I/O size of the service; matches the 128 KiB transfer of the NVMe-oF target.
[media], [raid], [data], [compression]—as deliveredWritten by the installer for this card and its SSDs.
[debug]—as deliveredInternal; do not change.

LMCache plugin

adapter_params of the native_plugin L2 adapter in lmcache server --l2-adapter '{…}' (LMCache L2 plugin). All keys are optional; an unknown key is refused with the list of accepted ones.

KeyDefaultWhat it does
port, gateway_host5557, localhostThe gateway.
fragment_bytes131072I/O size and fragment width: at least 4096 and at most the device’s transfer limit. Set it equal to the gateway’s fragment_size when that is set (261120 on a local card). Changing it makes everything stored a miss.
verifysampleoff, sample (first and last 4 KiB of every fragment) or full.
batch_slots, get_inflight, put_inflight, exist_batch_keys1024, 4, 2, 8192Request size in ring slots and pipelining. A request never takes more than half the ring.
copy_threads, copy_chunk_frags8, 32Copy and checksum workers, and their job size in fragments.
max_pending_store_mb16384Store backlog above which new stores are refused; 0 = no limit.
wait_timeout_s, reopen_after_s30, 30How long the gateway may stay silent before the session is given up; the delay before a new one (0 = never reopen).
namespaceemptyMixed into every key: separate deployments on one card. Changing it makes everything stored a miss.
start_cleanfalseWipe the card when the plugin opens — for tests only.
prefaulttrueTouch the ring pages at open.
stats_interval_s, stats_path, log_task_min_objects, log_level60, empty, 64, infoObservability: a summary line in the server log every stats_interval_s, JSON statistics to a file.

Leave LMCache’s own L2 limit max_capacity_gb at 0: space on the card is managed by the card’s key eviction.

OffloadingSpec backend

xkv.<key> entries (or a nested xkv object) in kv_connector_extra_config; the environment variable XKV_OFFLOAD_PARAMS, a JSON object, is merged last (OffloadingSpec backend).

Key (xkv. prefix)DefaultWhat it does
port5557The gateway.
fragment_bytes131072I/O and fragment size; 261120 for a local card, at most the transfer limit of a remote one.
staging_bytes, staging_min_slots2 GiB, 4The pool of staging slots.
pipeline_depth2Waves of transfers in flight at once.
max_objects_per_wave00 = the pool divided by pipeline_depth.
device_stagingautoauto: slots in GPU memory, one host copy; off: two copies through pinned host memory.
tp_layoutautoauto / replicated: one object per tensor-parallel group for MLA pages; per_rank: one per rank (with port_per_rank: true).
port_per_rankfalseA gateway per rank, for per_rank.
sync_lookup, sync_lookup_timeout_msfalse, 200Probe a new request’s keys inside the scheduler step instead of deferring the request by a step.
trust_own_storesfalseWhether the scheduler counts its own stores as present without asking the card.
probe_batch_keys, absent_ttl_s, max_index_entries4096, 60, 1000000The index of what the card holds.
op_timeout_s, acquire_timeout_s120, 30Timeouts of an operation and of acquiring a slot.
draft_groupsautoWhich KV groups of a hybrid model with MTP count as draft groups: auto keeps the Mamba groups out of them (needed for hits after a restart); vllm uses vLLM’s own marking.
flush_orderautoauto completes pending stores before model runner V2 reuses their blocks; vllm keeps vLLM’s order.
timingfalseA log line per wave and per job.

Settings that require a clean cache

The layout of what is stored is derived from the configuration and not recorded on the card. After any of these changes, wipe the card (how):

ChangeWhat happens to the old objects
fragment_size, fragment_spread_across_dbs (gateway)Not read back correctly (they are not reported as misses) — a wipe is mandatory.
slot_alignment_bytes (connector)The object size changes; old objects are unusable.
chunk_size, key_namespace (connector)Never found again; they only take space until evicted.
Tensor-parallel size, the model or its revisionNever found again.
A release whose notes say the stored layout changedAs the release notes say; wipe.

A change of PYTHONHASHSEED or of the vLLM version can change the block hashes as well; the cache then refills, and nothing wrong is ever read.

Planning a deployment?

Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.

Talk to Awide Labs