STORY · MODELLER_

vLLM 0.20.0 drastically improves efficiency – new open models launched

vLLM has released version 0.20.0 with significant improvements in memory and MoE serving, including TurboQuant 2-bit KV cache that delivers 4× increased capacity and 2.1% latency improvement. Meanwhile, Poolside is launching the code model Laguna XS.2 (33B/3B MoE) under Apache 2.0 license, and NVIDIA is presenting Nemotron 3 Nano Omni, a 30B multimodal MoE model with 256K context window.

WHY IT MATTERS

These developments make it significantly cheaper and faster to run large models locally, especially on modest GPUs, which accelerates the democratization of advanced AI. It also marks a shift away from NVIDIA's CUDA monopoly toward heterogeneous accelerator solutions.

SOURCES

MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.