← Back to the directory

University / Lab · Works with National University of Singapore, Tsinghua University

Tianjin University

天津大学


Coverage2

Research result · Sep 22, 2026

CUHK-Shenzhen-led team proposes COBRA-Skills, cutting agent skill optimization cost by 55%–58% with contextual bandits

The Chinese University of Hong Kong, Shenzhen, together with Tianjin University, The Hong Kong University of Science and Technology (Guangzhou) and National University of Singapore, details COBRA-Skills in a paper titled “COBRA-SKILLS: Contextual Bandit-Guided Evolution for Agent Skill Optimization.” The work models agent skill optimization as a budget-controlled contextual multi-armed bandit problem: a lightweight MLP predicts reward and LinearUCB quantifies exploration to choose candidate skills for real evaluation, while regeneration, trajectory mutation and crossover recombination operators evolve the population on a logarithmic schedule. The authors report that using only 50 optimization samples cuts optimization cost by 55%–58%, outperforming existing mainstream methods on six cross-domain benchmarks. On Qwen3.6-35B-A3B, GPT-5.4-Nano and Gemma-4-26B-A4B-it, improvements over no-skill baselines reach 13.1, 26.9 and 22.5 percentage points.

Original sources (Chinese)

对话 COBRA-Skills 一作卢平琛:仅用 50 个样本,重构 Agent 技能降本路线leiphone

Research result · Sep 15, 2026 · as partner

Yuanli Lingji's DM0.5 tops all four RoboColiseum sub-leaderboards

On the RoboColiseum benchmark built by Zhiyuan Robotics with universities and the open-source community, Yuanli Lingji's embodied foundation model DM0.5 ranks first on all four sub-leaderboards: instruction following (0.8444), spatial understanding (0.6146), disturbance adaptation (0.7344), and general manipulation (0.6370). The company says DM0.5 is currently the only model topping all four boards; it attributes the result to a strong pretraining base plus standard supervised fine-tuning, without task-specific optimization.

The model's core latency dropped from 534 ms to 57.49 ms, and it has native 60-second memory. In logistics sorting it averages about 3 seconds per item with accuracy above 99%, according to the company. Separately, GeoVLA, a collaboration between Yuanli Lingji, Tsinghua University and Tianjin University, was nominated for the IROS 2026 Cognitive Robotics Best Paper Award.

Original sources (Chinese)

横扫四榜,DM0.5 凭什么面面俱到?leiphone肉眼看不出的幻觉?清华提出视觉源幻觉,仅用0.9%数据实现SOTAaiera