Even GPT-5.2 Can't Count to Five: The Case for Zero-Error Horizons in Trustworthy LLMs

LLM HDFE LNJR AI Safety / AI Ethics AVAR
我们提出了“零错误视界”(Zero-Error Horizon,简称ZEH)这一概念,用于衡量大语言模型(LLM)的可信度,其定义为模型在不产生任何错误的前提下所能解决任务的最大范围。尽管ZEH本身形式简洁,但我们证明:对当前最先进大语言模型的ZEH进行评估,能够揭示大量富有启发性的洞见。例如,通过对GPT-5.2开展ZEH评估,我们发现:该模型甚至无法正确计算一个极短字符串(如“11000”)的奇偶性,也无法判断括号串“(((())))))”是否匹配平衡——而这一结果令人惊讶,毕竟GPT-5.2在其他方面展现出卓越的能力。大语言模型在如此基础的问题上仍会出错,这一事实为将其部署于安全关键型领域敲响了重要警钟。进一步地,我们将ZEH方法应用于Qwen2.5并展开细致分析,结果表明:虽然ZEH与模型整体准确率存在相关性,但其具体行为模式却各不相同;更重要的是,ZEH还能为算法能力(algorithmic capabilities)的涌现提供关键线索。最后,尽管ZEH的计算开销较大,我们亦探讨了若干优化策略,例如借助树状结构与在线Softmax技术,可将计算速度提升高达一个数量级,从而有效缓解该成本问题。
We propose Zero-Error Horizon (ZEH) for trustworthy LLMs, which represents the maximum range that a model can solve without any errors. While ZEH itself is simple, we demonstrate that evaluating the ZEH of state-of-the-art LLMs yields abundant insights. For example, by evaluating the ZEH of GPT-5.2, we found that GPT-5.2 cannot even compute the parity of a short string like 11000, and GPT-5.2 cannot determine whether the parentheses in ((((()))))) are balanced. This is surprising given the excellent capabilities of GPT-5.2. The fact that LLMs make mistakes on such simple problems serves as an important lesson when applying LLMs to safety-critical domains. By applying ZEH to Qwen2.5 and conducting detailed analysis, we found that while ZEH correlates with accuracy, the detailed behaviors differ, and ZEH provides clues about the emergence of algorithmic capabilities. Finally, while computing ZEH incurs significant computational cost, we discuss how to mitigate this cost by achieving up to one order of magnitude speedup using tree structures and online softmax.
许愿