vLLM v0.28.0 release adds major Kimi-K3 and DeepSeek V4 optimizations, speculative decoding upgrades, and broader ROCm support
EDITOR BRIEF
The vLLM project released v0.28.0 with 584 commits from 270 contributors, including broad performance work for Kimi-K3 and DeepSeek V4. Highlights include faster decode and prefill kernels, speculative decoding improvements, memory-saving shared-expert sharding, Model Runner V2 maturation, and expanded ROCm enablement.
INSIGHTS
The release shows vLLM continuing to evolve from an inference server into a high-performance optimization layer for frontier and open-weight models. Its focus on speculative decoding, ROCm, and memory efficiency reflects industry demand for lower latency, better GPU utilization, and more hardware-flexible AI deployment.
COMMENTS
Discussion
> geekhaus:~$ next read?

