AINEWS Search
Topics and information types

基础模型 研究

Back Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · Original publication time AI score54

Virginia Tech paper proposes compacting older-token caches in looped language modelsMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI introduction

Looped language models pass each token through the same layers multiple times, increasing cache use. The paper proposes keeping the latest 128 tokens exact and compressing older tokens into small vectors that each loop can read directly. On Ouro models at 16K tokens, throughput rose by up to 7.4 times, while accuracy on math, knowledge, and reasoning stayed above 97% of the original.

Article · Original

The article text is unavailable in this language; an existing version is shown.

New Virginia Tech paper shows how cache older tokens as compact vectors that attention reads directly, and a looped LLM fits 4.0 to 8.8 times as many requests per GPU at little accuracy cost.

Looped models run each token through the same layers several times, and every pass adds to the KV cache, so fewer requests fit on a GPU.

Hybrid Latent Attention keeps the last 128 tokens exact and squeezes older ones into small vectors that every loop reads without rebuilding anything.

On Ouro models, throughput rose up to 7.4 times at 16K tokens, and accuracy stayed above 97% of the original on math, knowledge and reasoning.

Gains are largest on long contexts. You can retrofit an existing looped checkpoint by training only the new parts.

– arxiv. org/abs/2610.07940

Title: "Hybrid Latent Attention for Looped Language Models"

来源:Rohan Paul · x.com

Research
Found an error? Send a correction