AINEWS 搜索
主题与信息类型

基础模型 研究

返回 Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · 原发布时间 AI 评分54

Virginia Tech 论文提出压缩循环语言模型的旧词元缓存

自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。

AI 导读

循环语言模型会让每个词元多次经过同一组网络层,增加缓存占用。论文提出保留最近 128 个词元的完整缓存,将更早的词元压缩为可供各轮直接读取的小向量。在 Ouro 模型的 16K 词元测试中,吞吐量最高提升 7.4 倍,数学、知识和推理任务的准确率保持在原模型的 97% 以上。

正文 · 原文

该语言的正文暂不可用,当前显示已有版本。

New Virginia Tech paper shows how cache older tokens as compact vectors that attention reads directly, and a looped LLM fits 4.0 to 8.8 times as many requests per GPU at little accuracy cost.

Looped models run each token through the same layers several times, and every pass adds to the KV cache, so fewer requests fit on a GPU.

Hybrid Latent Attention keeps the last 128 tokens exact and squeezes older ones into small vectors that every loop reads without rebuilding anything.

On Ouro models, throughput rose up to 7.4 times at 16K tokens, and accuracy stayed above 97% of the original on math, knowledge and reasoning.

Gains are largest on long contexts. You can retrofit an existing looped checkpoint by training only the new parts.

– arxiv. org/abs/2610.07940

Title: "Hybrid Latent Attention for Looped Language Models"

来源:Rohan Paul · x.com

论文
发现内容有误?提交纠错