← Back to the latest

ReleaseOct 5, 2026Entry № 857

VA-Bench released: a closed-loop Observe-Reason-Act-Revise benchmark for embodied spatial intelligence of multimodal models


VA-Bench Team has released VA-Bench, a benchmark for embodied spatial intelligence that scores multimodal models on a closed-loop Observe-Reason-Act-Revise cycle. It spans 14 robot manipulation task types—11 single-arm and 3 dual-arm—with 280 physically validated scenarios, and was used to evaluate 12 mainstream multimodal models.

On subtasks, target localization reached near 100% and manipulation semantics 99.6–100%, but complete-task success stayed around half: Qwen3.8-max at 53.93%, Opus-5 at 52.86%, and GPT-5.6-sol at 51.55%. The team identifies active observation, fine spatial execution, online error correction, and dual-arm coordination as the main bottlenecks; in ablation tests, removing active observation consistently lowered success rates.

Original sources (Chinese)

看懂不等于做对:VA-Bench测出大模型空间智能的执行断层jiqizhixin