AINEWS PLUS 2026-09-18 搜索
2026-09-18 UTC+8INDEPENDENT PERSPECTIVES.
← 返回动态
自动精选
综合 / 官方发布

Release: v5.16.0

来源摘要

# Release v5.16.0 ## New Model additions ### Qwen4-Exp Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE). GR is a Qwen-developed residual architecture that combines Hyper-Connection with GatedNorm. It mixes multiple residual streams with fine-grained elementwise gating before each attention and Mixture-of-Experts (MoE) block, then controls how much of the block output is injected back into each stream. QSA uses multiple query heads to score compressed key blocks, selects the most relevant contiguous token blocks, and k

关键事实

# Release v5.16.0 ## New Model additions ### Qwen4-Exp Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE). GR is a Qwen-developed residual architecture that combines Hyper-Connection with GatedNorm. It mixes multiple residual streams with fine-grained elementwise gating before each attention and Mixture-of-Experts (MoE) block, then controls how much of the block output is injected back into each stream. QSA uses multiple query heads to score compressed key blocks, selects the most relevant contiguous token blocks, and k

阅读原文

自动收录自 Transformers 发布记录 的公开订阅信息;请以原始出处为准。

本站不保存或展示该来源的全文。

前往原始出处阅读 ↗
来源:Transformers 发布记录 · 版本 6 · 最后更新 09月18日 09:25 · 新闻信息不构成投资建议。
发现内容有误?提交纠错