AINEWS PLUS 2026-09-18 搜索
2026-09-18 UTC+8INDEPENDENT PERSPECTIVES.
← 返回动态
自动精选
综合 / 官方发布

v0.25.0

来源摘要

# vLLM v0.25.0 Release Notes ## Highlights This release features 558 commits from 232 contributors (64 new)! * **Model Runner V2 is now the default for all dense models** (#44443). Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953). * **PagedAttention has been removed** (#47361). The legacy attention implementation is deleted now that V1/MRv2 backends ar

关键事实

# vLLM v0.25.0 Release Notes ## Highlights This release features 558 commits from 232 contributors (64 new)! * **Model Runner V2 is now the default for all dense models** (#44443). Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953). * **PagedAttention has been removed** (#47361). The legacy attention implementation is deleted now that V1/MRv2 backends ar

阅读原文

自动收录自 vLLM 发布记录 的公开订阅信息;请以原始出处为准。

本站不保存或展示该来源的全文。

前往原始出处阅读 ↗
来源:vLLM 发布记录 · 版本 6 · 最后更新 09月18日 09:25 · 新闻信息不构成投资建议。
发现内容有误?提交纠错