Inferact releases TPU inference kernel; Kimi K3 test outpaces GB200 under stated conditionsMachine translation
Inferact's megakernel combines Kimi K3's 92 layers into one Pallas program and runs with DSpark speculative decoding. In a test using 16 TPU v7 Ironwood chips versus 16 GB200 chips, both with vLLM, throughput was 709 versus 452 tokens per second. The kernel is currently tailored to Kimi K3 and needs adaptation for other model architectures.
Inferact releases TPU inference kernel; Kimi K3 test outpaces GB200 under stated conditions
Inferact's megakernel combines Kimi K3's 92 layers into one Pallas program and runs with DSpark speculative decoding. In a test using 16 TPU v7 Ironwood chips versus 16 GB200 chips, both with vLLM, throughput was 709 versus 452 tokens per second. The kernel is currently tailored to Kimi K3 and needs adaptation for other model architectures.
材料 d5efcef586d046b8b2c6274c2ff78070;建议 c89dab0a4a5a481494f53fb40d985ee5;系统证据核验通过,非人工审稿。
Read at the original source