Research result · Sep 24, 2026
METASTONE's Meta-Infer engine boosts DeepSeek inference throughput up to 6.87x on PCIe-only GPUs
In tests with its Meta-Infer inference engine, METASTONE (是石科技) reports raising DeepSeek-V4.1-Flash input throughput from a community baseline of 1,932 tok/s to 13,274 tok/s across eight PCIe-only GPUs, a 6.87x increase with 1M-token context. The company attributes the gain to kernel completion, operator optimization, communication restructuring, parallelism and capacity tuning, and cache reuse, without changing model structure or task semantics.
On GLM5.3, throughput rose 1.92x and P95 first-token latency fell from 141.6 seconds to 46.6 seconds; MiniMax H3 video generation was 2.48x faster end-to-end. METASTONE estimates roughly 1.5 machines of 6000D match one B300 reference image-to-video throughput, and says the same method has been validated on domestic GPUs.
Original sources (Chinese)
PCIe显卡被低估了!内核补齐+通信重构,DeepSeek推理吞吐翻近7倍