MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

AK's Picks LLM Other LLM Agent Multi-agent Collaboration Agent Eval Benchmarks
大型语言模型(LLMs)作为自主代理展示了非凡的能力,但现有的基准测试要么专注于单个代理任务,要么局限于狭窄的领域,无法捕捉多代理协调和竞争的动态。在本文中,我们介绍了MultiAgentBench,这是一个全面的基准测试框架,旨在评估基于LLM的多代理系统在多样化的互动场景中的表现。我们的框架不仅衡量任务完成情况,还使用新颖的、基于里程碑的关键绩效指标来评估协作和竞争的质量。此外,我们评估了各种协调协议(包括星型、链型、树型和图结构拓扑)以及创新策略,如小组讨论和认知规划。值得注意的是,gpt-4o-mini 在任务得分上达到了平均最高分,在研究场景中,图结构在协调协议中表现最佳,而认知规划使里程碑达成率提高了3%。代码和数据集可在 https://github.com/MultiagentBench/MARBLE 公开获取。
Large Language Models (LLMs) have shown remarkable capabilities as autonomous agents, yet existing benchmarks either focus on single-agent tasks or are confined to narrow domains, failing to capture the dynamics of multi-agent coordination and competition. In this paper, we introduce MultiAgentBench, a comprehensive benchmark designed to evaluate LLM-based multi-agent systems across diverse, interactive scenarios. Our framework measures not only task completion but also the quality of collaboration and competition using novel, milestone-based key performance indicators. Moreover, we evaluate various coordination protocols (including star, chain, tree, and graph topologies) and innovative strategies such as group discussion and cognitive planning. Notably, gpt-4o-mini reaches the average highest task score, graph structure performs the best among coordination protocols in the research scenario, and cognitive planning improves milestone achievement rates by 3%. Code and datasets are public available at https://github.com/MultiagentBench/MARBLE.
许愿