A Survey on LLM-as-a-Judge

LLM HDCE Other LLM NLP Benchmarks Other NLP
准确和一致的评估对于众多领域的决策至关重要,但由于固有的主观性、变化性和规模问题,这仍然是一项具有挑战性的任务。大型语言模型(LLMs)在多个领域取得了显著的成功,这导致了“LLM作为法官”这一概念的出现,即使用LLM作为复杂任务的评估者。由于能够处理多种数据类型并提供可扩展、成本效益高且一致的评估,LLM成为传统专家驱动评估的一种有吸引力的替代方案。然而,确保“LLM作为法官”系统的可靠性仍然是一个重大挑战,需要精心设计和标准化。本文对“LLM作为法官”进行了全面的综述,探讨了核心问题:如何构建可靠的“LLM作为法官”系统?我们探索了提高可靠性的策略,包括提高一致性、减少偏见以及适应多样化的评估场景。此外,我们提出了评估“LLM作为法官”系统可靠性的方法,并为此设计了一个新的基准。为了推动“LLM作为法官”系统的开发和实际应用,我们还讨论了实际应用、挑战和未来方向。本综述为这一快速发展的领域的研究人员和实践者提供了基础参考。
Accurate and consistent evaluation is crucial for decision-making across numerous fields, yet it remains a challenging task due to inherent subjectivity, variability, and scale. Large Language Models (LLMs) have achieved remarkable success across diverse domains, leading to the emergence of "LLM-as-a-Judge," where LLMs are employed as evaluators for complex tasks. With their ability to process diverse data types and provide scalable, cost-effective, and consistent assessments, LLMs present a compelling alternative to traditional expert-driven evaluations. However, ensuring the reliability of LLM-as-a-Judge systems remains a significant challenge that requires careful design and standardization. This paper provides a comprehensive survey of LLM-as-a-Judge, addressing the core question: How can reliable LLM-as-a-Judge systems be built? We explore strategies to enhance reliability, including improving consistency, mitigating biases, and adapting to diverse assessment scenarios. Additionally, we propose methodologies for evaluating the reliability of LLM-as-a-Judge systems, supported by a novel benchmark designed for this purpose. To advance the development and real-world deployment of LLM-as-a-Judge systems, we also discussed practical applications, challenges, and future directions. This survey serves as a foundational reference for researchers and practitioners in this rapidly evolving field.
许愿