Research result · Oct 7, 2026
HIT Shenzhen and NUS propose QuantWM: training-free 2-bit KV Cache quantization for video world models with up to 6.20x compression
QuantWM is a training-free 2-bit KV cache quantization framework for video world models from Miao Zhang's group at Harbin Institute of Technology, Shenzhen (iLearn-Lab) and Shuicheng Yan's group at National University of Singapore, with Jiaqi Zhao and Xiaobin Hu. Existing 2-bit methods can score near BF16 yet generate flicker, blur and detail loss because Key quantization error changes attention-score rankings, causing wrong historical frame or spatial position choices. QuantWM uses quantization sensitivity-aware clustering (QSAC) and principal subspace attention compensation (PSAC). Across Matrix-Game-2, LingBot-World-v2, HY-World 1.5, LongCat-Video and Causal-Forcing, the authors report better PSNR, SSIM and LPIPS than QVG and KIVI, a reduction in the highest historical token selection change ratio from 53.88% to 10.19%, and up to 6.20x KV cache compression using a custom Triton kernel.
Original sources (Chinese)
量化后画面闪烁?时序稳定的视频世界模型2-bit KV Cache压缩方案诞生