Release · Sep 11, 2026
China Mobile Cloud and partners release China's first domestic GPU + neuromorphic chip heterogeneous hybrid LLM inference system
At the 2026 China Computing Power Conference, China Mobile Cloud, CETC Nanhu Research Institute, Lynxi Technologies, Iluvatar CoreX, Tsinghua University and Peking University released what they describe as China’s first domestic GPU + neuromorphic chip heterogeneous hybrid LLM inference system. The system splits large-model computation: attention work goes to domestic GPUs while latency-sensitive FFN/MoE expert modules run on neuromorphic chips, coordinated by a self-developed compiler, interconnect protocol and unified inference engine.
In tests with DeepSeek V4, the partners say the setup improves cost-performance more than twofold compared with similar domestic GPU clusters and cuts operating costs by over 40%. It is aimed at token factories, AI code generation, multi-agent collaboration and smart manufacturing.
Original sources (Chinese)
国内首个国产 GPU + 类脑芯片大模型异构混合推理系统发布,较同类国产 GPU 算力集群性价比提升一倍以上