GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

LLM LNJR Other LLM
最近大型语言模型(LLM)的进展引起了人们对它们的形式推理能力的兴趣,特别是在数学方面。GSM8K基准广泛用于评估模型在小学级别问题上的数学推理能力。虽然LLM在GSM8K上的表现在最近几年中显着提高,但它们的数学推理能力是否真正提高仍不清楚,这引发了对报告指标可靠性的质疑。为了解决这些问题,我们对几个SOTA开放和封闭模型进行了大规模研究。为了克服现有评估的局限性,我们引入了GSM-Symbolic,这是一个基于符号模板的改进基准,可以生成各种各样的问题。GSM-Symbolic能够进行更可控的评估,提供关键见解和更可靠的度量,以衡量模型的推理能力。我们的研究发现,LLM在回答同一问题的不同实例时表现出明显的差异。具体而言,当只改变GSM-Symbolic基准中问题中的数值时,所有模型的性能都会下降。此外,我们研究了这些模型中数学推理的脆弱性,并表明随着问题子句数量的增加,它们的性能显着下降。我们假设这种下降是因为当前的LLM无法进行真正的逻辑推理;它们复制训练数据中的推理步骤。即使一个似乎与问题相关的子句不对最终答案所需的推理链做出贡献,添加一个单子句也会导致所有最先进的模型的性能显着下降(高达65%)。总的来说,我们的工作提供了对LLM在数学推理方面能力和局限性的更细致的理解。
Recent advancements in Large Language Models (LLMs) have sparked interest in their formal reasoning capabilities, particularly in mathematics. The GSM8K benchmark is widely used to assess the mathematical reasoning of models on grade-school-level questions. While the performance of LLMs on GSM8K has significantly improved in recent years, it remains unclear whether their mathematical reasoning capabilities have genuinely advanced, raising questions about the reliability of the reported metrics. To address these concerns, we conduct a large-scale study on several SOTA open and closed models. To overcome the limitations of existing evaluations, we introduce GSM-Symbolic, an improved benchmark created from symbolic templates that allow for the generation of a diverse set of questions. GSM-Symbolic enables more controllable evaluations, providing key insights and more reliable metrics for measuring the reasoning capabilities of models.Our findings reveal that LLMs exhibit noticeable variance when responding to different instantiations of the same question. Specifically, the performance of all models declines when only the numerical values in the question are altered in the GSM-Symbolic benchmark. Furthermore, we investigate the fragility of mathematical reasoning in these models and show that their performance significantly deteriorates as the number of clauses in a question increases. We hypothesize that this decline is because current LLMs cannot perform genuine logical reasoning; they replicate reasoning steps from their training data. Adding a single clause that seems relevant to the question causes significant performance drops (up to 65%) across all state-of-the-art models, even though the clause doesn't contribute to the reasoning chain needed for the final answer. Overall, our work offers a more nuanced understanding of LLMs' capabilities and limitations in mathematical reasoning.
许愿