PAPER

Revisiting Prompt Optimization with Large Reasoning Models-A Case Study on Event Extraction

LLM CoT NLP NER APGO
2025年04月10日
大型推理模型(LRMs),例如 DeepSeek-R1 和 OpenAI 的 o1,在各种推理任务中展现了卓越的能力。它们在生成和处理中间推理步骤方面的强大能力,也引发了这样的讨论:这些模型可能不再需要大量的提示工程或优化,即可正确理解人类指令并生成准确的输出。在本研究中,我们旨在系统性地探讨这一开放性问题,并以事件抽取这一结构化任务为案例进行研究。我们对两种 LRMs(DeepSeek-R1 和 o1)以及两种通用型大语言模型(LLMs,GPT-4o 和 GPT-4.5)进行了实验,测试了它们作为任务模型或提示优化器时的表现。结果表明,在像事件抽取这样复杂的任务中,LRMs 作为任务模型时仍然可以从提示优化中获益;而将 LRMs 用作提示优化器时,则可以生成更有效的提示。最后,我们对 LRMs 常见的错误进行了分析,并强调了 LRMs 在改进任务指令和事件指南时表现出的稳定性和一致性。
Large Reasoning Models (LRMs) such as DeepSeek-R1 and OpenAI o1 have demonstrated remarkable capabilities in various reasoning tasks. Their strong capability to generate and reason over intermediate thoughts has also led to arguments that they may no longer require extensive prompt engineering or optimization to interpret human instructions and produce accurate outputs. In this work, we aim to systematically study this open question, using the structured task of event extraction for a case study. We experimented with two LRMs (DeepSeek-R1 and o1) and two general-purpose Large Language Models (LLMs) (GPT-4o and GPT-4.5), when they were used as task models or prompt optimizers. Our results show that on tasks as complicated as event extraction, LRMs as task models still benefit from prompt optimization, and that using LRMs as prompt optimizers yields more effective prompts. Finally, we provide an error analysis of common errors made by LRMs and highlight the stability and consistency of LRMs in refining task instructions and event guidelines.
许愿