← Back to the directory

University / Lab · Works with OpenAgents, Columbia University

University of Pennsylvania

宾夕法尼亚大学


Coverage1

Research result · Oct 10, 2026 · as partner

OpenAgents and universities release AgentWorld, a multi-agent collaboration benchmark, with top model hitting only 52% success

OpenAgents, with researchers from Columbia University, University of Pennsylvania, Seoul National University and Pennsylvania State University, released AgentWorld, a multi-agent collaboration benchmark built in an MMORPG sandbox. It uses 100 human-designed tasks and 100 augmented variants requiring 3 to 20 agents, with asymmetric roles, black-box interactions and long-horizon execution exposed through 13 high-level API tools. On the main task set, Gemini 3 Flash reached 52.0% success, Claude Haiku 4.5 45.0%, GPT-5 Mini 36.0%, and DeepSeek R1-70B 20.0%; the lowest augmented-task score was 10.0%. The paper also introduces CCE, a causal collaboration effectiveness metric that traces actions contributing to success.

Original sources (Chinese)

Agent组队干活,最强模型只完成50%任务!基准评测协作能力aiera