← Back to the directory

Individual · Works with National University of Singapore, Harbin Institute of Technology, Shenzhen (iLearn-Lab)

Shuicheng Yan

颜水成


Coverage1

Research result · Oct 7, 2026 · as partner

HIT Shenzhen and NUS propose QuantWM: training-free 2-bit KV Cache quantization for video world models with up to 6.20x compression

QuantWM is a training-free 2-bit KV cache quantization framework for video world models from Miao Zhang's group at Harbin Institute of Technology, Shenzhen (iLearn-Lab) and Shuicheng Yan's group at National University of Singapore, with Jiaqi Zhao and Xiaobin Hu. Existing 2-bit methods can score near BF16 yet generate flicker, blur and detail loss because Key quantization error changes attention-score rankings, causing wrong historical frame or spatial position choices. QuantWM uses quantization sensitivity-aware clustering (QSAC) and principal subspace attention compensation (PSAC). Across Matrix-Game-2, LingBot-World-v2, HY-World 1.5, LongCat-Video and Causal-Forcing, the authors report better PSNR, SSIM and LPIPS than QVG and KIVI, a reduction in the highest historical token selection change ratio from 53.88% to 10.19%, and up to 6.20x KV cache compression using a custom Triton kernel.

Original sources (Chinese)

量化后画面闪烁?时序稳定的视频世界模型2-bit KV Cache压缩方案诞生jiqizhixin