A common way to speed up KV cache restores is to keep the cache in host RAM, or to put a DRAM tier in front of storage. We measured what that buys over the ADP KV cache connector alone, which restores straight from NVMe drives through the accelerator and keeps no cache in RAM. The two came out roughly level. LMCache holding the entire working set in 400 GiB of DRAM returned a 196,000-token prompt on a running engine 0.18 seconds (10%) sooner than pure ADP, which uses no expensive DRAM. That was its only faster restore: after an engine restart its DRAM was empty and ADP was faster. With concurrent agents both generated at the same rate, and at 32 agents LMCache had the longer tail. The difference is the bill: ADP gets the same speed without the RAM, which makes it substantially cheaper to run.
Is a KV cache in RAM faster than one on NVMe behind the ADP accelerator?
Restoring a long prompt from a KV cache instead of recomputing it is what makes long-context agents affordable. On GLM-5.3 a 196,000-token prompt takes 47.7 seconds to prefill from scratch and about two seconds to bring back. The intuition is that the medium decides the rest: keep the cache in host RAM, or at least put a RAM tier in front of the drives, and restores get faster still. LMCache in its multiprocess (MP) mode is the standard way to build that. A separate cache server keeps an L1 in DRAM and can pass misses to an L2 below it.
We tested the intuition head-on. LMCache 0.5.5 got an L1 of 400 GiB of pinned DRAM, more than the whole working set of the test (about 230 GB), with a retain policy so nothing was evicted. We compared it with the ADP connector alone on the same hardware, model and prompts. First as a pure RAM cache, then as the top of a two-tier cache with ADP below it.
| Setup | Where a restored page comes from |
|---|---|
| ADP connector alone | the accelerator and its NVMe drives, over NVMe-oF; no RAM tier |
| LMCache, DRAM only | 400 GiB of host DRAM; nothing below it |
| LMCache, DRAM + ADP | 400 GiB of host DRAM first, the ADP accelerator on a miss: the classic multi-tier setup |
RAM restores 0.18 s (10%) faster on a warm engine, and nowhere else
First, a single 196,000-token prompt, restored two ways. With the engine running and only its GPU prefix cache cleared, every setup still holds the prompt. After a restart of the vLLM server the GPU memory is empty and the prompt has to come back from wherever it survived.
On a running engine, the prompt comes back from 400 GiB of RAM in 1.68 s and from the NVMe drives in 1.86 s: RAM is ahead by 0.18 s, 10% of the connector's time. With the same RAM in front of ADP, the multi-tier arrangement, the gap shrinks to 0.11 s, within the spread of the multi-tier runs themselves (1.63–2.07 s). This is the only restore RAM wins: after a restart the connector is faster, and under load, below, returning turns come back from the drives as fast or faster.
After a vLLM restart, only the drives still hold the prompt. The connector brings it back in 2.26 s. In this deployment the LMCache server runs in the engine's container, so its DRAM restarts with the engine and comes back empty. With ADP below it, the prompt comes back through LMCache in 2.91 s. We saw the empty L1 directly: after the restart LMCache read all 10.75 GB of the prompt from ADP. With RAM alone there is no lower tier to read from, so the prompt is recomputed; we did not time that case separately. A RAM cache meant to survive restarts has to run as a separate, long-lived service; we did not test that arrangement.
The drive is not the bottleneck
A restore of this prompt moves about 10.7 GB of KV pages into eight GPUs and then prefills the few hundred tokens past the last cached block. Neither step gets faster when the pages sit in host RAM: they cross the same PCIe links into GPU memory either way, and the tail is compute.
What the connector does is keep the drives from ever being the slow part. The gateway keeps dozens of commands in flight and spreads the fragments of every object across the accelerator's databases, so the accelerator reads from its NVMe drives in parallel. The connector loads the cache layer by layer, so layer n is already computing while layer n+1 is still arriving. Each GPU copies only its share of a layer and NVLink fans the rest out. The same parallelism runs on the write side, so storing a prompt does not hold up generation. The result is a restore within 0.2 seconds of RAM, from drives that keep their contents through any restart.
With concurrent agents, the two are level, and RAM has the longer tail
A single warm restore is the best case for RAM. Production traffic is many agents at once, each returning to a long conversation, while new turns are written to the cache. We replayed anonymized coding-agent traces at 8, 16 and 32 concurrent agents against the ADP connector alone and against LMCache with its 400 GiB DRAM tier and ADP below it. The DRAM tier holds the whole working set and evicts nothing, so under load LMCache serves from RAM.
The decode rates are close. ADP trails by 14% at 8 agents and by 7% at 16, and the two meet at 32. The price of that small lead is 400 GiB of pinned DRAM on every node. Both find the same share of returning turns in the external cache: 72, 88 and 90% for the connector, 74, 88 and 90% for LMCache. At full load, serving the pages from RAM bought no lead in generation.
The difference is in the tail, and it grows with the load. At 32 agents it is hard to miss.
| Metric | ADP connector alone | LMCache, DRAM + ADP |
|---|---|---|
| Decode rate per agent, tokens/s | 21.0 / 27.3 / 27.0 | 24.5 / 29.2 / 27.2 |
| Hits from the external cache | 72 / 88 / 90% | 74 / 88 / 90% |
| First token of a returning turn, median, s | 9.2 / 37.2 / 73.9 | 16.6 / 36.8 / 111.2 |
| Answers frozen for more than 20 s | 0 / 0 / 0 | 0 / 0 / 27 |
| Wall time of the step, s | 346 / 458 / 661 | 367 / 417 / 915 |
At 32 agents a returning turn waits a median 111 seconds for its first token with LMCache and 74 with the connector. Of its 380 answers, 27 froze mid-generation for more than 20 seconds, 24 of them for more than a minute; with the connector, no answer paused for longer than 5.7 seconds. The whole step takes 915 seconds against 661: ADP finishes it 28% sooner. At 8 and 16 agents the picture is mixed: LMCache finishes the 16-agent step sooner, the connector brings returning turns back sooner at 8.
The same speed, without the RAM
If the speed is about the same, the question is what each setup takes from the node. The connector keeps no cache in RAM at all: the only host memory it uses is a 10 GiB ring through which pages pass on their way between the drives and the GPUs. LMCache holds its cache in 400 GiB of pinned DRAM, and the cache server's resident size reached 412 GiB.
- RAM is the expensive medium. Server DRAM costs many times more per gigabyte than NVMe flash, and 400 GiB of it per node is sized for one model's working set. A bigger working set, or a second model, means more RAM on every node.
- The storage is needed anyway. A RAM cache empties at every engine restart and holds only as much as the RAM, so a cache that has to survive restarts needs drives below it. A RAM tier is paid on top of that storage, not instead of it.
- Another service to run. The cache server has to be sized, monitored and upgraded alongside vLLM, and to keep its contents through engine restarts it has to live outside the engine's container.
- A second tier to keep right. Multi-tier caching means two places that can disagree and a write path that has to be correct in both.
For all of that, the gain is 0.18 seconds, 10%, on a warm, single restore of a 196,000-token prompt.
Level on speed, far cheaper on memory: deploy the connector
For serving long-context agents from vLLM, LMCache with its KV cache in 400 GiB of host RAM and the ADP KV cache connector reading from NVMe drives are roughly level. RAM is 10% faster on a warm restore, the connector is faster after a restart, and under concurrent agents both generate at the same rate and find the same share of turns in the cache, with a longer tail for LMCache at 32 agents.
What separates them is the memory. The connector needs no RAM tier: no hundreds of gigabytes of pinned DRAM per node, no separate cache service, and its cache survives restarts on the drives. For the same speed, that makes it substantially cheaper to deploy and to scale. There is no need to buy RAM or build multi-tier memory for the KV cache: deploy the connector.
Test setup
- Hardware and model
- GLM-5.3 FP8 on 8× NVIDIA H200, TP=8, vLLM 0.28.0 with prefix caching; the ADP accelerator reached over NVMe-oF.
- LMCache
- LMCache 0.5.5 in MP mode, L1 of 400 GiB pinned DRAM with a retain policy, larger than the working set of the test (about 230 GB); the cache server in the vLLM container. In the two-tier setup, our L2 plugin for ADP below it.
- Single restore
- A 196,000-token prompt; after a reset of the GPU prefix cache (three times) and after a restart of the vLLM server (three times for the connector, once for LMCache with ADP). Time to first token at the client.
- Under load
- Anonymized coding-agent traces at 8, 16 and 32 concurrent agents; the ADP connector alone against LMCache with DRAM and ADP. One run per setup, on the same hardware over two days, with a few unrelated requests in some of the windows; no noise band, so differences of a few percent are within spread.
Planning memory for long-context inference?
Before sizing a RAM tier, measure the restore you would get without one. We are happy to run the comparison on your model.