Sparse Layers are Critical to Scaling Looped Language Models

LLM MQIA MoE
循环式语言模型通过在深度方向上重复使用一组Transformer层,从而降低了内存开销,并在每次循环的边界处自然地提供了早期退出(early-exit)点。然而,与采用独立层的标准Transformer相比,循环式模型在扩展性方面表现欠佳。我们对比了标准Transformer与混合专家(Mixture-of-Experts, MoE)Transformer,分别考察其带循环与不带循环两种架构,得出两项主要发现。 第一,我们发现循环式MoE模型(Looped-MoE)的扩展性优于标准基线模型,而循环式稠密模型(dense looped model)则无法做到这一点。这一差异源于循环之间的路由发散(routing divergence):在循环式MoE模型中,同一组共享层在每次循环遍历时会激活不同的专家,从而在不增加参数量的前提下恢复了模型的表达能力。 第二,我们发现循环式模型配合早期退出机制,在计算开销与生成质量之间的权衡上优于标准模型。这是因为每一次循环结束时所经过的层,恰好就是最终输出所依赖的相同层,因此循环边界天然成为更优的退出点;实证结果也证实,在这些边界点上,模型输出更早趋于收敛。 综上所述,本研究为循环式模型的规模化发展指明了一条清晰路径:采用“循环式MoE + 早期退出”的架构,不仅能在大规模场景下超越标准Transformer,还能在几乎不损失生成质量的前提下,显著节省内存占用与推理计算开销。
Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundaries. However, looped models do not scale as favorably as standard transformers with unique layers. We compare standard and Mixture-of-Experts (MoE) transformers, with and without looping, and find two main results. First, we find Looped-MoE models scale better than the standard baseline while dense looped models do not. We trace this to routing divergence between loops: in Looped-MoE models, different experts are activated on each pass through the same shared layers, recovering expressivity without additional parameters. Our second finding is that looped models have better compute-quality trade-offs with early exits than standard models. Because each loop ends with the same layers that produce the final output, loop boundaries are superior exit points, as confirmed by earlier output convergence at these points. In sum, we provide a clear direction for scaling looped models: a Looped-MoE model with early exits can not only beat standard transformers at scale, but also enable significant memory and inference savings with minimal degradation in quality.
许愿