GEEK HAUS
Back to feed

vLLM v0.28.0 release adds major Kimi-K3 and DeepSeek V4 optimizations, speculative decoding upgrades, and broader ROCm support

·github.com
read original

EDITOR BRIEF

The vLLM project released v0.28.0 with 584 commits from 270 contributors, including broad performance work for Kimi-K3 and DeepSeek V4. Highlights include faster decode and prefill kernels, speculative decoding improvements, memory-saving shared-expert sharding, Model Runner V2 maturation, and expanded ROCm enablement.

INSIGHTS

The release shows vLLM continuing to evolve from an inference server into a high-performance optimization layer for frontier and open-weight models. Its focus on speculative decoding, ROCm, and memory efficiency reflects industry demand for lower latency, better GPU utilization, and more hardware-flexible AI deployment.

COMMENTS

Discussion

> geekhaus:~$ next read?

Next read recommendations