Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization

LLM RLHF ML Offline RL
2024年10月11日
强化学习(RL)在将大型语言模型(LLMs)与人类偏好对齐以及提高它们执行复杂任务的能力方面起着至关重要的作用。然而,当前的方法要么由于使用多个模型和广泛的在线采样进行训练而需要大量的计算资源(例如PPO),要么被构建为赌博问题(例如DPO,DRO),这些方法往往难以处理多步推理任务,例如数学问题解决和涉及长时间思考的复杂推理。为了克服这些限制,我们介绍了直接Q函数优化(DQO),它将响应生成过程构造为马尔可夫决策过程(MDP),并利用软演员-评论家(SAC)框架直接优化由语言模型参数化的Q函数。 DQO的MDP公式比基于赌博的方法具有结构上的优势,可以更有效地进行过程监督。在两个数学问题解决数据集GSM8K和MATH上的实验结果表明,DQO优于先前的方法,将其确立为一种有前途的离线强化学习方法,用于对齐语言模型。
Reinforcement Learning (RL) plays a crucial role in aligning large language models (LLMs) with human preferences and improving their ability to perform complex tasks. However, current approaches either require significant computational resources due to the use of multiple models and extensive online sampling for training (e.g., PPO) or are framed as bandit problems (e.g., DPO, DRO), which often struggle with multi-step reasoning tasks, such as math problem-solving and complex reasoning that involve long chains of thought. To overcome these limitations, we introduce Direct Q-function Optimization (DQO), which formulates the response generation process as a Markov Decision Process (MDP) and utilizes the soft actor-critic (SAC) framework to optimize a Q-function directly parameterized by the language model. The MDP formulation of DQO offers structural advantages over bandit-based methods, enabling more effective process supervision. Experimental results on two math problem-solving datasets, GSM8K and MATH, demonstrate that DQO outperforms previous methods, establishing it as a promising offline reinforcement learning approach for aligning language models.
许愿