GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

LLM Latent Reasoning CoT NLP Benchmarks Agent Human Agent Interaction Agent Eval Benchmarks
我们推出了GDPval,这是一个用于评估人工智能模型在现实世界中具有经济价值任务上表现能力的基准。GDPval涵盖了美国劳工统计局所列44个职业中的大部分工作活动,这些职业来自对美国国内生产总值(GDP)贡献最大的前九大行业。这些任务基于拥有平均14年经验的行业专业人士的实际工作内容构建而成。我们发现,前沿模型在GDPval上的表现随时间大致呈线性提升,当前最先进的模型在交付成果质量方面已接近行业专家水平。我们分析了前沿模型在辅以人类监督的情况下,完成GDPval任务的成本和速度相较于无辅助的人类专家是否更具优势。此外,我们还证明,增加推理投入、扩充任务上下文信息以及加强任务结构化支持均能提升模型在GDPval上的表现。最后,我们开源了一个包含220项任务的高质量子集,并在evals.openai.com提供公开的自动化评分服务,以促进未来对模型现实世界能力的研究。
We introduce GDPval, a benchmark evaluating AI model capabilities on real-world economically valuable tasks. GDPval covers the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDP (Gross Domestic Product). Tasks are constructed from the representative work of industry professionals with an average of 14 years of experience. We find that frontier model performance on GDPval is improving roughly linearly over time, and that the current best frontier models are approaching industry experts in deliverable quality. We analyze the potential for frontier models, when paired with human oversight, to perform GDPval tasks cheaper and faster than unaided experts. We also demonstrate that increased reasoning effort, increased task context, and increased scaffolding improves model performance on GDPval. Finally, we open-source a gold subset of 220 tasks and provide a public automated grading service at evals.openai.com to facilitate future research in understanding real-world model capabilities.
许愿