Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

AK's Picks LLM RLHF Latent Reasoning ML RL PGPO
并行思维作为一种新方法逐渐兴起,旨在通过同时探索多条推理路径来增强大语言模型(LLMs)的推理能力。然而,通过训练激活这种能力仍面临挑战,因为现有方法主要依赖于在合成数据上进行监督式微调(SFT),这种方式更倾向于强制模型模仿教师模型,而非鼓励探索与泛化能力。与这些方法不同,我们提出了\textbf{Parallel-R1},这是首个面向复杂现实世界推理任务的并行思维强化学习(RL)框架。 我们的框架采用了一种渐进式的课程设计,明确解决了使用强化学习训练并行思维时遇到的冷启动问题。我们首先在较简单任务生成的提示轨迹上进行监督式微调,以培养模型的并行思维能力;随后过渡到强化学习阶段,使模型能够在更复杂的问题上进行探索和泛化。在多个数学基准测试(包括MATH、AMC23和AIME)上的实验表明,Parallel-R1 成功培养了模型的并行思维能力,在准确率上相比仅通过强化学习直接在复杂任务上训练的顺序思维模型提升了8.4%。 进一步分析显示,模型的思维行为发生了明显转变:在训练早期阶段,并行思维被用作一种探索策略;而在后期阶段,该能力则被用于从多个视角进行验证。更重要的是,我们将并行思维验证为一种\textbf{训练中期的探索支架},这种短暂的探索阶段为后续强化学习带来了更高的性能上限,在AIME25上的表现相比基线提升了42.9%。我们的模型、数据和代码将开源,地址为 https://github.com/zhengkid/Parallel-R1。
Parallel thinking has emerged as a novel approach for enhancing the reasoning capabilities of large language models (LLMs) by exploring multiple reasoning paths concurrently. However, activating such capabilities through training remains challenging, as existing methods predominantly rely on supervised fine-tuning (SFT) over synthetic data, which encourages teacher-forced imitation rather than exploration and generalization. Different from them, we propose \textbf{Parallel-R1}, the first reinforcement learning (RL) framework that enables parallel thinking behaviors for complex real-world reasoning tasks. Our framework employs a progressive curriculum that explicitly addresses the cold-start problem in training parallel thinking with RL. We first use SFT on prompt-generated trajectories from easier tasks to instill the parallel thinking ability, then transition to RL to explore and generalize this skill on harder problems. Experiments on various math benchmarks, including MATH, AMC23, and AIME, show that Parallel-R1 successfully instills parallel thinking, leading to 8.4% accuracy improvements over the sequential thinking model trained directly on challenging tasks with RL. Further analysis reveals a clear shift in the model's thinking behavior: at an early stage, it uses parallel thinking as an exploration strategy, while in a later stage, it uses the same capability for multi-perspective verification. Most significantly, we validate parallel thinking as a \textbf{mid-training exploration scaffold}, where this temporary exploratory phase unlocks a higher performance ceiling after RL, yielding a 42.9% improvement over the baseline on AIME25. Our model, data, and code will be open-source at https://github.com/zhengkid/Parallel-R1.