Engineering notes, benchmarks, release reports, and company news from the team building the Awide Data Processor.
Release v6.1.1 of the ADP KV cache connector carries per-tenant cache isolation into the external KV store, loads model weights about six times faster, and keeps generation running when the storage gateway goes away. Verified with a full regression on hardware and already serving production traffic.
We gave LMCache 400 GiB of DRAM, enough for the whole working set, and measured it against the ADP KV cache connector alone, which restores from NVMe drives and keeps no cache in RAM. The two came out roughly level: RAM was 10% faster on a warm restore, the connector faster after a restart, the same decode rate under load. The difference is the 400 GiB of RAM per node that the connector does not need.
Prefix caching shares computed KV blocks between every request that starts the same way, and the time to first token tells anyone whether a prompt was already cached. How cache isolation closes that channel, what vLLM gives you out of the box, what the API front end has to do, and when isolation is worth its cost.
MiniMax M3 on vLLM 0.28 verified on hardware with a full regression, support for the Kimi K3 architecture on vLLM 0.29, NVLink fan-out of cache restores, and the first steps toward a connector that new models no longer break.
A new release brings hardware-verified support for GLM-5.3 and DeepSeek-V4-Flash on vLLM 0.28/0.29, deployment tooling, and improved stability for context-heavy inference workloads.
A new release of the ADP KV cache connector: vLLM 0.26, plus the model families that used to break KV offload — hybrid Mamba (HMA) and multi-token prediction (MTP).
Awide Labs has acquired the intellectual property license and manufacturing rights for the Pliops technology portfolio, bringing hardware-accelerated data processing and key-value technology together with its PostgreSQL expertise and AI roadmap.