Autodata: An agentic data scientist to create high quality synthetic data

Agent Agent Planning Autonomous Workflows ML ML
我们提出了“Autodata”——一种通用方法,使人工智能代理能够充当数据科学家,自主构建高质量的训练与评估数据。我们展示了如何训练(即元优化)此类数据科学家代理,使其习得生成更优质数据的能力。文中阐述了该方法的整体建模框架,并给出了一个具体、实用的实现方案——“代理式自指示”(Agentic Self-Instruct)。我们在计算机科学研究任务、法律推理任务以及涉及数学对象的推理任务上开展了实验,结果表明,相较于传统的合成数据集构建方法,本方法均取得了更优性能。此外,对数据科学家代理本身进行元优化,还能带来更为显著的性能提升。这种基于代理的数据生成范式,提供了一条将更多推理算力转化为更高品质模型训练数据的有效路径。总体而言,我们认为这一研究方向有望从根本上改变人工智能数据的构建方式。
We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data. We show how to train (meta-optimize) such a data scientist agent, so that it learns to create even stronger data. We describe the overall formulation, and a specific practical implementation, Agentic Self-Instruct. We conduct experiments on computer science research tasks, legal reasoning tasks and reasoning with mathematical objects, where we obtain improved results compared to classical synthetic dataset creation methods. Further, meta-optimizing the data scientist agent itself delivers an even larger performance uplift. Agentic data creation provides a way to convert increased inference compute into higher quality model training. Overall, we believe this direction has the potential to change the way we build AI data.
许愿