Research resultOct 10, 2026Entry № 937
OpenAgents and universities release AgentWorld, a multi-agent collaboration benchmark, with top model hitting only 52% success
OpenAgents, with researchers from Columbia University, University of Pennsylvania, Seoul National University and Pennsylvania State University, released AgentWorld, a multi-agent collaboration benchmark built in an MMORPG sandbox. It uses 100 human-designed tasks and 100 augmented variants requiring 3 to 20 agents, with asymmetric roles, black-box interactions and long-horizon execution exposed through 13 high-level API tools. On the main task set, Gemini 3 Flash reached 52.0% success, Claude Haiku 4.5 45.0%, GPT-5 Mini 36.0%, and DeepSeek R1-70B 20.0%; the lowest augmented-task score was 10.0%. The paper also introduces CCE, a causal collaboration effectiveness metric that traces actions contributing to success.
Original sources (Chinese)
Agent组队干活,最强模型只完成50%任务!基准评测协作能力