PaperBanana: Automating Academic Illustration for AI Scientists

AK's Picks Multimodal Intelligence Vision-Language Pre-training Agent Agent Planning Autonomous Workflows GenAI TDIG Inpainting Other GenAI
尽管依托大语言模型的自主式人工智能科研助手发展迅猛,但生成符合出版要求的学术插图仍是当前研究工作流中一项费时费力的瓶颈任务。为减轻这一负担,我们提出了 PaperBanana——一个面向自动化生成出版级学术插图的智能体框架。PaperBanana 依托当前最先进的视觉语言模型(VLM)与图像生成模型,协调多个专业化智能体协同工作,依次完成参考文献检索、内容与风格规划、图像渲染,以及通过自反思机制进行迭代优化。为对本框架开展严格评估,我们构建了 PaperBananaBench 评测基准,其中包含从 NeurIPS 2025 会议论文中精心筛选出的 292 个方法示意图测试用例,覆盖广泛的研究领域与多样的插图风格。大量实验结果表明,PaperBanana 在忠实性、简洁性、可读性与美学质量等各项指标上均持续优于现有主流基线方法。我们还进一步验证了该方法可有效拓展至高质量统计图表的生成任务。总体而言,PaperBanana 为出版级学术插图的自动化生成开辟了新路径。
Despite rapid advances in autonomous AI scientists powered by language models, generating publication-ready illustrations remains a labor-intensive bottleneck in the research workflow. To lift this burden, we introduce PaperBanana, an agentic framework for automated generation of publication-ready academic illustrations. Powered by state-of-the-art VLMs and image generation models, PaperBanana orchestrates specialized agents to retrieve references, plan content and style, render images, and iteratively refine via self-critique. To rigorously evaluate our framework, we introduce PaperBananaBench, comprising 292 test cases for methodology diagrams curated from NeurIPS 2025 publications, covering diverse research domains and illustration styles. Comprehensive experiments demonstrate that PaperBanana consistently outperforms leading baselines in faithfulness, conciseness, readability, and aesthetics. We further show that our method effectively extends to the generation of high-quality statistical plots. Collectively, PaperBanana paves the way for the automated generation of publication-ready illustrations.
许愿