AINEWS PLUS 2026-09-18 搜索
2026-09-18 UTC+8INDEPENDENT PERSPECTIVES.
返回
精选

v0.28.0

vLLM 发布记录··原发布时间
内容导读

# v0.28.0 ## Highlights This release features 584 commits from 270 contributors (76 new)! * **Kimi-K3 performance push**: a major optimization effort for Kimi-K3 across the stack — Decode Context Parallel (DCP) support (#50484), fused FlashKDA decode and prefill kernels (#50654, #51311, #52458), SiTU activation support for MegaMoE (#50510), GEMM-RS for sequence parallelism (#52079), combined all-gathers with 1.5~3x kernel-level speedup (#51070), an adaptive speculative token budget delivering ~60% better DSpark TTFT (#51725), and optional shared-expert sharding saving ~17 GiB of memory per GPU (#50912). Kimi-K3 also now runs on ROCm with the

推荐理由

来自 vLLM 发布记录 的官方公开更新;本站仅提供短摘要与原始出处链接。

v0.28.0

· 原发布时间
内容导读

# v0.28.0 ## Highlights This release features 584 commits from 270 contributors (76 new)! * **Kimi-K3 performance push**: a major optimization effort for Kimi-K3 across the stack — Decode Context Parallel (DCP) support (#50484), fused FlashKDA decode and prefill kernels (#50654, #51311, #52458), SiTU activation support for MegaMoE (#50510), GEMM-RS for sequence parallelism (#52079), combined all-gathers with 1.5~3x kernel-level speedup (#51070), an adaptive speculative token budget delivering ~60% better DSpark TTFT (#51725), and optional shared-expert sharding saving ~17 GiB of memory per GPU (#50912). Kimi-K3 also now runs on ROCm with the

推荐理由

来自 vLLM 发布记录 的官方公开更新;本站仅提供短摘要与原始出处链接。

自动收录自 vLLM 发布记录 的公开订阅信息;请以原始出处为准。

前往原始出处阅读
发现内容有误?提交纠错