Human vs. Machine: Behavioral Differences Between Expert Humans and Language Models in Wargame Simulations

LLM Other LLM Agent Multi-agent Collaboration Human Agent Interaction Agent Eval Benchmarks AI Safety / AI Ethics EVA RLAW
有些人认为,人工智能(AI)的出现可以提供更好的决策和增加军事效力,同时减少人为错误和情感的影响。然而,关于AI系统,特别是大型语言模型(LLMs)在高风险军事决策情境中与人类的表现相比如何,仍存在争议。这种情况可能会增加升级和不必要冲突的风险。为了测试这种潜在情况并审查LLMs在此类目的中的应用,我们使用了一个新的战争游戏实验,涉及107名国家安全专家,旨在研究虚构的美中情境下的危机升级,并将人类玩家与LLM模拟响应进行了分别模拟的比较。战争游戏在军事战略的发展和国家对威胁或攻击的反应方面具有悠久的历史。在这里,我们展示了LLM和人类响应的相当高水平的一致性,以及个体行动和战略倾向的显著数量和质量差异。这些差异取决于LLMs对于遵循战略指令后适当的暴力水平的内在偏见、LLM的选择以及LLMs是否被要求直接为玩家团队做出决策或首先模拟玩家之间的对话。当模拟对话时,讨论缺乏质量并保持荒谬的和谐。LLM模拟无法考虑到人类玩家的特征,即使是极端特征,如“和平主义者”或“侵略性社交病态者”,也没有显著差异。我们的结果促使决策者在授予自主权或遵循基于AI的战略建议之前要谨慎。
To some, the advent of artificial intelligence (AI) promises better decision-making and increased military effectiveness while reducing the influence of human error and emotions. However, there is still debate about how AI systems, especially large language models (LLMs), behave compared to humans in high-stakes military decision-making scenarios with the potential for increased risks towards escalation and unnecessary conflicts. To test this potential and scrutinize the use of LLMs for such purposes, we use a new wargame experiment with 107 national security experts designed to look at crisis escalation in a fictional US-China scenario and compare human players to LLM-simulated responses in separate simulations. Wargames have a long history in the development of military strategy and the response of nations to threats or attacks. Here, we show a considerable high-level agreement in the LLM and human responses and significant quantitative and qualitative differences in individual actions and strategic tendencies. These differences depend on intrinsic biases in LLMs regarding the appropriate level of violence following strategic instructions, the choice of LLM, and whether the LLMs are tasked to decide for a team of players directly or first to simulate dialog between players. When simulating the dialog, the discussions lack quality and maintain a farcical harmony. The LLM simulations cannot account for human player characteristics, showing no significant difference even for extreme traits, such as "pacifist" or "aggressive sociopath". Our results motivate policymakers to be cautious before granting autonomy or following AI-based strategy recommendations.
许愿