PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing

NLP TSIC Benchmarks Agent Multi-agent Collaboration Agent Eval Benchmarks
将非结构化的研究资料整合撰写成学术论文,是人工智能驱动科学发现过程中一项至关重要却尚未得到充分探索的挑战。现有的自主写作系统往往与特定实验流程僵化绑定,所生成的文献综述也流于表面、缺乏深度。为此,我们提出“PaperOrchestra”——一种面向自动化AI研究论文撰写的多智能体框架。该框架能够灵活地将任意形式的前期写作素材(即无约束的原始材料)转化为可直接投稿的LaTeX格式论文,不仅涵盖全面、深入的文献综述,还自动生成各类可视化内容,例如数据图表与概念示意图。为客观评估系统性能,我们构建了“PaperWritingBench”,这是首个基于200篇顶级人工智能会议论文反向重构而成的标准化基准数据集,其中包含真实、完整的原始写作材料;同时配套开发了一整套自动化评测工具。在并排式人工评估中,“PaperOrchestra”显著优于现有各类自主写作基线方法:在文献综述质量方面,其绝对胜率高出基线模型50%–68%;在整篇论文质量方面,绝对胜率优势亦达14%–38%。
Synthesizing unstructured research materials into manuscripts is an essential yet under-explored challenge in AI-driven scientific discovery. Existing autonomous writers are rigidly coupled to specific experimental pipelines, and produce superficial literature reviews. We introduce PaperOrchestra, a multi-agent framework for automated AI research paper writing. It flexibly transforms unconstrained pre-writing materials into submission-ready LaTeX manuscripts, including comprehensive literature synthesis and generated visuals, such as plots and conceptual diagrams. To evaluate performance, we present PaperWritingBench, the first standardized benchmark of reverse-engineered raw materials from 200 top-tier AI conference papers, alongside a comprehensive suite of automated evaluators. In side-by-side human evaluations, PaperOrchestra significantly outperforms autonomous baselines, achieving an absolute win rate margin of 50%-68% in literature review quality, and 14%-38% in overall manuscript quality.