Documentation menu
Integration options

LMCache L2 plugin

Supported

The ADP card as the L2 tier of the LMCache multiprocess server: how it fits under LMCache, how to run it, its settings, and what to expect from it.

What it is

The plugin makes the ADP card the second tier (L2) of the LMCache multiprocess server. LMCache keeps everything it does — its interface to vLLM, its index, its policies and its CPU tier (L1, pinned host memory); objects that leave L1, or are looked up after a restart, are stored on the card behind the storage gateway. LMCache loads the plugin through its own native_plugin adapter type: LMCache itself is not changed.

The ADP card under LMCache
LMCache unchanged · the plugin is its L2
vLLM LMCacheMPConnector in every rank LMCACHE SERVER · UNCHANGED L1 · CPU tier pinned host memory index, eviction policy L2 adapter type native_plugin loads the ADP plugin ADP L2 plugin C++, no Python lock Storage gateway fragments, databases ADP card NVMe SSDs behind it batches ring + control port STORE objects leaving L1 go to the card LOAD L1 hit: host memory otherwise, and after a restart: card → L1 → GPU

Use it when LMCache is already part of your stack. If it is not, the vLLM KV connector restores faster and needs no host-memory tier (comparison).

How it stores data

  • Objects and fragments. One LMCache chunk is one object; it is cut into fragments of fragment_bytes, each one I/O. A small manifest is written only after every fragment of the object has been acknowledged: its presence is what “the object is on the card” means.
  • Integrity. A load checks the manifest — size, geometry, key — and, by default, a sample checksum of every fragment. Anything that does not match is a miss for that object only, and an unusable manifest is removed so the next store writes the object again.
  • Never written twice. Before a store the plugin asks the card whether the object is already there, and skips it if so: a key is a hash of the tokens, so the same chunk is the same entry.
  • Failures. A missing key or fragment is a miss. A gateway that stops answering (wait_timeout_s) turns every answer into a miss until the plugin opens a new session (reopen_after_s) — which is also how a restarted gateway is picked up. A gateway that is not up yet when the server starts is the same case.
  • Backpressure. Stores beyond max_pending_store_mb of backlog are refused at once, so L1 objects are not held up behind a slow card; loads take priority over stores.

Requirements

vLLM0.28.0
LMCache0.5.4, as bundled with the vllm/vllm-openai:v0.28.0 image, in multiprocess (server) mode
ImagevLLM 0.28.0 with the plugin compiled on it, delivered by Awide Labs on request
Gateway and cardas for the connector: ADP card setup, the gateway from Installation

Run it

The LMCache server takes the plugin as its L2 adapter:

lmcache server --host 127.0.0.1 --port 6555 --l1-size-gb 120 --eviction-policy LRU --chunk-size 256 \
  --l2-adapter '{"type":"native_plugin","module_path":"xkv_lmcache","class_name":"XkvL2Connector",
                 "adapter_params":{"port":5557,"fragment_bytes":131072}}'

and vLLM connects to the LMCache server:

--kv-transfer-config '{"kv_connector":"LMCacheMPConnector",
  "kv_connector_module_path":"lmcache.integration.vllm.lmcache_mp_connector",
  "kv_role":"kv_both","kv_connector_extra_config":{"lmcache.mp.port":6555}}'

Inside the vLLM container. The image starts the LMCache server from the container’s main process when LMC_ENABLE=1 is set, and waits until its port answers. The server and its L1 then restart with vLLM.

LMC_ENABLE=1 LMC_L1_GB=120 LMC_CHUNK=256 LMC_XKV_PORT=5557 LMC_XKV_FRAG=131072
# optional: LMC_XKV_VERIFY, LMC_XKV_COPY_THREADS, LMC_XKV_STATS_PATH=/tmp/xkv-stats.json,
#           LMC_XKV_PARAMS='{"batch_slots":2048}' (merged last); python3 -m xkv_lmcache_autostart --help lists them all

As a separate process, so that L1 survives a restart of vLLM: the same image, started with python3 -m xkv_lmcache_autostart --exec (without LMC_ENABLE). It needs the GPUs (it copies L1 to the GPU through CUDA IPC), the host IPC namespace and /dev/shm shared with vLLM and the gateway, and the gateway port:

docker run -d --name lmcache-server --gpus all --network host --ipc host \
  -e LMC_HOST=0.0.0.0 -e LMC_PORT=6555 -e LMC_L1_GB=120 -e LMC_CHUNK=256 \
  -e LMC_XKV_PORT=5557 -e LMC_XKV_HOST=<gateway address> -e LMC_XKV_FRAG=131072 \
  --entrypoint python3 <image> -m xkv_lmcache_autostart --exec
# vLLM: "kv_connector_extra_config": {"lmcache.mp.host": "<server address>", "lmcache.mp.port": 6555}

On Kubernetes that is a second container in the model pod, or a host-network, host-IPC pod on the node.

Settings

The plugin is configured through the adapter_params of the L2 adapter (or the LMC_XKV_* variables of the autostart); every key, with its default, is in Configuration → LMCache plugin. Set fragment_bytes to the gateway’s fragment_size. Leave LMCache’s own L2 limit max_capacity_gb at 0: space on the card is managed by its key eviction.

What to expect

  • Hits in L1 come back slightly faster than from the card through the connector, at the cost of the host memory L1 takes. In our comparison on GLM-5.3, a 196,000-token prompt held in a 400 GiB L1 returned in 1.68 s against 1.86 s through the connector (details).
  • Hits from the card go card → L1 → GPU: LMCache finishes the whole prefetch into L1 before it copies to the GPU, so a restore after an engine restart is slower than through the connector (2.91 s against 2.26 s in the same comparison). That is inside LMCache, not the plugin.
  • Host memory: the data passes through two host buffers, the gateway ring and L1.

Limits: one LMCache server per node and one gateway per server; two servers must not write one card; x86_64 Linux only.

Monitoring

Every stats_interval_s the plugin writes a summary line prefixed [xkv-lmcache] to the LMCache server log; stats_path writes the full JSON: per-operation tasks, hits, bytes, latency percentiles, queue depths, ring use, integrity counters, session state and process memory. The store outcomes are told apart by store_skipped (already on the card), store_deduped, store_dropped (backlog full) and store_rejected_by_card — the last one growing means the card refuses writes (capacity).

Planning a deployment?

Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.

Talk to Awide Labs