← Back to the directory

University / Lab · Works with Harbin Institute of Technology, Shenzhen (iLearn-Lab), Miao Zhang

National University of Singapore

新加坡国立大学


Coverage2

Research result · Oct 7, 2026 · as partner

HIT Shenzhen and NUS propose QuantWM: training-free 2-bit KV Cache quantization for video world models with up to 6.20x compression

QuantWM is a training-free 2-bit KV cache quantization framework for video world models from Miao Zhang's group at Harbin Institute of Technology, Shenzhen (iLearn-Lab) and Shuicheng Yan's group at National University of Singapore, with Jiaqi Zhao and Xiaobin Hu. Existing 2-bit methods can score near BF16 yet generate flicker, blur and detail loss because Key quantization error changes attention-score rankings, causing wrong historical frame or spatial position choices. QuantWM uses quantization sensitivity-aware clustering (QSAC) and principal subspace attention compensation (PSAC). Across Matrix-Game-2, LingBot-World-v2, HY-World 1.5, LongCat-Video and Causal-Forcing, the authors report better PSNR, SSIM and LPIPS than QVG and KIVI, a reduction in the highest historical token selection change ratio from 53.88% to 10.19%, and up to 6.20x KV cache compression using a custom Triton kernel.

Original sources (Chinese)

量化后画面闪烁?时序稳定的视频世界模型2-bit KV Cache压缩方案诞生jiqizhixin

Research result · Sep 22, 2026 · as partner

CUHK-Shenzhen-led team proposes COBRA-Skills, cutting agent skill optimization cost by 55%–58% with contextual bandits

The Chinese University of Hong Kong, Shenzhen, together with Tianjin University, The Hong Kong University of Science and Technology (Guangzhou) and National University of Singapore, details COBRA-Skills in a paper titled “COBRA-SKILLS: Contextual Bandit-Guided Evolution for Agent Skill Optimization.” The work models agent skill optimization as a budget-controlled contextual multi-armed bandit problem: a lightweight MLP predicts reward and LinearUCB quantifies exploration to choose candidate skills for real evaluation, while regeneration, trajectory mutation and crossover recombination operators evolve the population on a logarithmic schedule. The authors report that using only 50 optimization samples cuts optimization cost by 55%–58%, outperforming existing mainstream methods on six cross-domain benchmarks. On Qwen3.6-35B-A3B, GPT-5.4-Nano and Gemma-4-26B-A4B-it, improvements over no-skill baselines reach 13.1, 26.9 and 22.5 percentage points.

Original sources (Chinese)

对话 COBRA-Skills 一作卢平琛:仅用 50 个样本,重构 Agent 技能降本路线leiphone