AINEWS 2026-09-26 中/EN Search
2026-09-26 UTC+8
INDEPENDENT PERSPECTIVES.中/EN

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

Back
Back

Inferact releases TPU inference kernel; Kimi K3 test outpaces GB200 under stated conditionsMachine translation

量子位 AI 报道··Original publication time
AI-assisted summary

Inferact's megakernel combines Kimi K3's 92 layers into one Pallas program and runs with DSpark speculative decoding. In a test using 16 TPU v7 Ironwood chips versus 16 GB200 chips, both with vLLM, throughput was 709 versus 452 tokens per second. The kernel is currently tailored to Kimi K3 and needs adaptation for other model architectures.

Inferact releases TPU inference kernel; Kimi K3 test outpaces GB200 under stated conditions

· 原发布时间
AI-assisted summary

Inferact's megakernel combines Kimi K3's 92 layers into one Pallas program and runs with DSpark speculative decoding. In a test using 16 TPU v7 Ironwood chips versus 16 GB200 chips, both with vLLM, throughput was 709 versus 452 tokens per second. The kernel is currently tailored to Kimi K3 and needs adaptation for other model architectures.

材料 d5efcef586d046b8b2c6274c2ff78070;建议 c89dab0a4a5a481494f53fb40d985ee5;系统证据核验通过,非人工审稿。

Read at the original source
发现内容有误?提交纠错