ReleaseSep 15, 2026Entry № 442
Fudan University CVL Lab Releases AVTrack Benchmark for Audio-Visual Speaker Tracking, Accepted at ICML 2026
A new audio-visual speaker tracking benchmark, AVTrack, is now public from Fudan University Computer Vision Lab, with the paper accepted at ICML 2026. It contains 871 videos averaging 54.0 seconds, yielding 3,120 pixel-level instance trajectories with cross-frame identities, and organizes eight challenge categories including visual occlusion, camera motion changes, multi-instance scenes, multi-round speech, and audio-visual inconsistency. The lab also releases AVTracker, a training-free modular method that reports 29.08 HOTA in unified evaluation, ahead of the existing AVIS method. Code, dataset, and paper are available.
Original sources (Chinese)
复旦发布以人为中心的复杂场景音视频追踪基准 | ICML 2026In the directoryFudan University Computer Vision Lab复旦大学计算机视觉实验室 · 1 entry