Verified models
Every model below passes the same regression on hardware: each KV page restored from the ADP card is compared byte for byte with what the engine wrote.
| Model | Attention architecture | vLLM | MTP |
|---|---|---|---|
| DeepSeek-V4-Flash | Compressed sparse MLA + sliding window | 0.29.0 | — |
| Kimi K3 | Kimi Delta Attention + MLA | 0.29.0 | — |
| Kimi-Linear-48B-A3B | Kimi Delta Attention + MLA | 0.29.0 | — |
| GLM-5.3 FP8 | Sparse MLA (DSA) | 0.28.0 | — |
| MiniMax M3 MXFP8 | Sparse attention with indexer, MoE | 0.28.0 | — |
| GLM-5.2 FP8 | Sparse MLA (DSA) | 0.26 | yes |
| GLM-4.7-Flash | MLA, AWQ | 0.26 | — |
| DeepSeek-V2-Lite | MLA | 0.26 | — |
| Gemma 4 12B | Full + sliding-window attention | 0.26 | yes |
| Qwen 3.8 27B | Linear + full attention | 0.26 | yes |
| Qwen 3.6 35B-A3B | Gated DeltaNet, MoE | 0.26 | yes |
| Qwen 3.6 27B | Mamba hybrid, three cache groups | 0.26 | — |
| Nemotron 3 Nano 30B-A3B | Mamba-2 + sparse attention | 0.26 | — |
The vLLM column is the version each model is verified on; one connector code base serves vLLM 0.26, 0.28 and 0.29.
IMPORTANT
Hybrid models (Mamba, Mamba-2, Gated DeltaNet): use vLLM 0.28.0 or newer with the connector. From 0.28 on, the vLLM scheduler aligns a partial prefix hit of a hybrid model with a stored recurrent state; vLLM 0.26 does not do this when a KV connector is configured. The connector keeps the stored cache consistent on every version.
Other models of the same architectures usually work as they are — the connector decides at start how every cache of the model is exchanged, and refuses to start, naming the cache, when it cannot. Talk to us about a model that is not on the list.
What each architecture needs
| Architecture | Examples | What the deployment needs |
|---|---|---|
| Full attention, MLA | DeepSeek-V2-Lite, GLM-4.7-Flash | chunk_size a multiple of the block size. |
| Full + sliding-window attention | Gemma 4 | The window at least twice chunk_size and a multiple of it. With deferred write (the default): --no-async-scheduling and pipeline parallelism 1, or set enable_deferred_write: false. The connector refuses to start otherwise and says which. |
| Mamba, Mamba-2, Gated DeltaNet, Kimi Delta Attention hybrids | Qwen 3.6, Qwen 3.8, Nemotron 3 Nano, Kimi-Linear, Kimi K3 | --enable-prefix-caching; --mamba-cache-mode align (vLLM 0.29 sets it by itself with prefix caching); --watermark=0.02 (why); pipeline parallelism 1. |
| DeepSeek Sparse Attention | GLM-5.2, GLM-5.3 | --kv-cache-dtype fp8; on vLLM 0.29 FULL_DECODE_ONLY CUDA graphs. Needs a GPU vLLM’s sparse MLA kernels support (Hopper or Blackwell). The indexer’s cache is restored and written with its layer. |
| Compressed sparse MLA | DeepSeek-V4-Flash | Deferred write left on (the default). |
| Sparse attention with indexer | MiniMax M3 | --kv-cache-dtype fp8, to keep one page size per group. |
Multi-token prediction (MTP) works with every architecture that has an MTP head: the drafter’s layers are recognised and kept out of the exchange.
Settings per model
The connector keys are in kv_connector_extra_config; the vLLM options go on the vLLM command line next to the model’s
own (tool-call and reasoning parsers, context length).
| Model | Connector | vLLM |
|---|---|---|
| DeepSeek-V4-Flash | chunk_size: 512 | vLLM 0.29.0, VLLM_USE_V2_MODEL_RUNNER=0. The first requests can pause for 15–23 s while DeepGEMM compiles kernels: warm up with representative requests and keep VLLM_CACHE_ROOT on a persistent volume. |
| Kimi K3 | chunk_size a multiple of the attention block vLLM sets (768 tokens at TP=8, 1536 with an fp8 KV cache); gateway fragment_size required — the page is 864 KiB at TP=8 | vLLM 0.29.0, VLLM_USE_V2_MODEL_RUNNER=0, --additional-config '{"kda_prefill_backend": "triton"}', --watermark=0.02, no --speculative-config |
| Kimi-Linear-48B-A3B | chunk_size a multiple of 256 at TP=8; gateway fragment_size required — the page is 288 KiB at TP=8 | vLLM 0.29.0, VLLM_USE_V2_MODEL_RUNNER=0, --watermark=0.02 |
| GLM-5.3 FP8 | chunk_size: 128; remote card: slot_alignment_bytes: 4096 | vLLM 0.28.0, --kv-cache-dtype fp8, --max-num-batched-tokens=8192. vLLM 0.28.0 is the recommended version for GLM-5.x: 0.29 decodes these models markedly slower, with or without the connector. |
| MiniMax M3 MXFP8 | — | vLLM 0.28.0, --kv-cache-dtype fp8 |
| GLM-5.2 FP8 | chunk_size: 128; remote card: slot_alignment_bytes: 4096 | vLLM 0.26 or 0.28.0, --kv-cache-dtype fp8, --max-num-batched-tokens=8192; MTP: '{"method":"mtp","num_speculative_tokens":1}' |
| GLM-4.7-Flash | chunk_size: 256 | — |
| DeepSeek-V2-Lite | — | — |
| Gemma 4 12B | chunk_size such that the sliding window is at least twice it and a multiple of it | --no-async-scheduling; MTP: '{"method":"gemma4_mtp","model":"google/gemma-4-12B-it-assistant","num_speculative_tokens":1}' |
| Qwen 3.8 27B | — (vLLM picks the block) | --mamba-cache-mode align, --watermark=0.02; MTP: '{"method":"qwen3_5_mtp","num_speculative_tokens":1}' |
| Qwen 3.6 35B-A3B | — | --mamba-cache-mode align, --watermark=0.02; MTP: '{"method":"qwen3_next_mtp","num_speculative_tokens":2}' |
| Qwen 3.6 27B | — | --mamba-cache-mode align, --watermark=0.02 |
| Nemotron 3 Nano 30B-A3B | — | --mamba-cache-mode align (required: this model would choose another mode by itself), --watermark=0.02 |
MTP values go to --speculative-config, in single quotes on a container command line. A dash in the connector column
means the defaults are right: chunk_size then follows vLLM’s block size, which vLLM raises for hybrid models so that
a page holds one recurrent state.
Choosing chunk_size
chunk_size is the number of tokens in one stored object. Larger chunks mean fewer and larger I/Os per restore; smaller
chunks a finer hit: a prompt of N tokens can hit at most floor(N / chunk_size) × chunk_size tokens, and a prompt
shorter than one chunk is never stored. It must be a multiple of the block size of every KV cache group (the connector
checks it at start). Keep the value from the table unless you have measured another one, and remember that changing it
starts a new cache.
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.