← Back to the latest

ReleaseSep 15, 2026Entry № 442

Fudan University CVL Lab Releases AVTrack Benchmark for Audio-Visual Speaker Tracking, Accepted at ICML 2026


A new audio-visual speaker tracking benchmark, AVTrack, is now public from Fudan University Computer Vision Lab, with the paper accepted at ICML 2026. It contains 871 videos averaging 54.0 seconds, yielding 3,120 pixel-level instance trajectories with cross-frame identities, and organizes eight challenge categories including visual occlusion, camera motion changes, multi-instance scenes, multi-round speech, and audio-visual inconsistency. The lab also releases AVTracker, a training-free modular method that reports 29.08 HOTA in unified evaluation, ahead of the existing AVIS method. Code, dataset, and paper are available.

Original sources (Chinese)

复旦发布以人为中心的复杂场景音视频追踪基准 | ICML 2026aiera