PAPER

Looped Diffusion Transformer

GenAI Diffusion TDIG ML TAVR
提升文本到图像模型的传统方法主要依赖于扩大模型规模或增加去噪步数。本文探索了一种替代性的计算扩展路径:在每个去噪步骤内部,反复运行共享的Transformer模块,从而在保持参数量不变的前提下,有效提升计算深度。这种“循环式计算”机制支持对模型内部表征进行迭代式精化,且无需显式引入推理用的特殊标记(reasoning tokens)。然而,简单地将模块循环执行并不能稳定地提升图像质量。我们发现,这一问题源于中间循环层所获得的监督信号较弱,以及注意力机制更新缺乏约束,导致局部细节信息在循环过程中逐步退化。为应对上述挑战,我们提出了“循环式扩散Transformer”(Looped-DiT),该方法将面向中间循环层的深层监督与自调制注意力机制(self-modulating attention)相结合,以稳定循环过程中的特征更新。在参数量与计算量均严格对齐的实验设置下,Looped-DiT始终显著优于非循环结构的基线模型。尤为值得注意的是,一个仅含2.6亿参数的循环模型,在多个文本到图像生成基准测试中,性能超越了一个参数量高达其6.5倍的更大模型,同时推理所需计算量却仅为后者的1/4.9。除性能优势外,我们还发现,循环式计算可为扩散模型提供一种更高效的迭代计算范式:在固定推理预算下,增加循环深度所带来的性能增益,明显高于单纯增加去噪步数;此外,更深的循环能够逐步修正前序循环中产生的错误,展现出潜在的隐式推理行为。综上所述,这些结果表明,循环式计算为视觉生成模型的规模化发展提供了一条极具前景的新路径。
Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit reasoning tokens. However, naive looping fails to consistently improve image quality. We trace this problem to weak supervision across intermediate loops and unregulated attention updates that progressively erode local information. To overcome these challenges, we propose Looped Diffusion Transformer (Looped-DiT), which combines deep supervision across intermediate loops with self-modulating attention to stabilize looped feature updates. Under matched-parameter and matched-compute settings, Looped-DiT consistently outperforms non-looped baselines. Notably, a 260M-parameter looped model can surpass a model 6.5x larger across multiple text-to-image benchmarks while requiring 4.9x lower inference compute. Beyond this performance gain, we find that looped computation can offer a more effective form of iterative computation for diffusion models, with increasing loop depth yielding larger gains than adding more denoising steps under a fixed inference budget. Furthermore, deeper loops can progressively correct mistakes made in earlier loops, exhibiting behaviors suggestive of latent reasoning. Together, these results show that looped computation offers a promising way to scale visual generation models.
许愿