METASTONE reports inference gains from its Meta-Infer engineMachine translation
METASTONE says Meta-Infer used kernel fixes, communication changes and parallelism tuning to raise DeepSeek-V4.1-Flash input throughput on eight PCIe-only GPUs from a community Day0 baseline of 1,932 tok/s to 13,274 tok/s, a 6.87-fold increase. The work did not change the model weights or structure; the figures come from tests reported in the article.
METASTONE reports inference gains from its Meta-Infer engine
METASTONE says Meta-Infer used kernel fixes, communication changes and parallelism tuning to raise DeepSeek-V4.1-Flash input throughput on eight PCIe-only GPUs from a community Day0 baseline of 1,932 tok/s to 13,274 tok/s, a 6.87-fold increase. The work did not change the model weights or structure; the figures come from tests reported in the article.
材料 733fc89b277c4f7099fbf9c601f2b1d9;建议 b3489a53428844128f9adbe12f3fa522;系统证据核验通过,非人工审稿。
Read at the original source