v0.22.1
## Highlights This release features 8 commits from 6 contributors (1 new)! v0.22.1 is a patch release on top of v0.22.0 with targeted bug fixes plus a couple of additions: new model support for JetBrains' Mellum v2, zentorch-accelerated quantized linear inference on AMD Zen CPUs, and fixes for multi-node Ray data-parallel serving, DeepSeek-V4 initialization, and a few model-loading regressions. ### Model Support * New model: JetBrains' **Mellum v2**, an open-weights Mixture-of-Experts code-generation model (#43992). * **DeepSeek-V4**: resolve a CUTLASS `fmin` compatibility issue that broke initialization (0decac0d). * Fix `OlmoHybridForCausal
来自 vLLM 发布记录 的官方公开更新;本站仅提供短摘要与原始出处链接。
v0.22.1
## Highlights This release features 8 commits from 6 contributors (1 new)! v0.22.1 is a patch release on top of v0.22.0 with targeted bug fixes plus a couple of additions: new model support for JetBrains' Mellum v2, zentorch-accelerated quantized linear inference on AMD Zen CPUs, and fixes for multi-node Ray data-parallel serving, DeepSeek-V4 initialization, and a few model-loading regressions. ### Model Support * New model: JetBrains' **Mellum v2**, an open-weights Mixture-of-Experts code-generation model (#43992). * **DeepSeek-V4**: resolve a CUTLASS `fmin` compatibility issue that broke initialization (0decac0d). * Fix `OlmoHybridForCausal
来自 vLLM 发布记录 的官方公开更新;本站仅提供短摘要与原始出处链接。
自动收录自 vLLM 发布记录 的公开订阅信息;请以原始出处为准。
前往原始出处阅读