← Back to the directory

University / Lab · Works with National University of Singapore

The Chinese University of Hong Kong, Shenzhen

香港中文大学(深圳)


Coverage1

Research result · Sep 22, 2026

CUHK-Shenzhen-led team proposes COBRA-Skills, cutting agent skill optimization cost by 55%–58% with contextual bandits

The Chinese University of Hong Kong, Shenzhen, together with Tianjin University, The Hong Kong University of Science and Technology (Guangzhou) and National University of Singapore, details COBRA-Skills in a paper titled “COBRA-SKILLS: Contextual Bandit-Guided Evolution for Agent Skill Optimization.” The work models agent skill optimization as a budget-controlled contextual multi-armed bandit problem: a lightweight MLP predicts reward and LinearUCB quantifies exploration to choose candidate skills for real evaluation, while regeneration, trajectory mutation and crossover recombination operators evolve the population on a logarithmic schedule. The authors report that using only 50 optimization samples cuts optimization cost by 55%–58%, outperforming existing mainstream methods on six cross-domain benchmarks. On Qwen3.6-35B-A3B, GPT-5.4-Nano and Gemma-4-26B-A4B-it, improvements over no-skill baselines reach 13.1, 26.9 and 22.5 percentage points.

Original sources (Chinese)

对话 COBRA-Skills 一作卢平琛:仅用 50 个样本,重构 Agent 技能降本路线leiphone