Discrete Flow Matching

LLM Latent Reasoning CoT ML Generative Models DM
2024年07月22日
尽管流匹配和扩散模型已成为生成连续变量(如图像和视频)的强大生成范式,但它们在生成高维离散数据(如语言)方面的应用仍然有限。在本文中,我们提出了离散流匹配(Discrete Flow Matching),这是一种专门设计用于生成离散数据的新型离散流范式。离散流匹配提供了几个关键贡献:(i)它可以使用在源分布和目标分布之间插值的一般概率路径族;(ii)它允许使用学习后验概率(如概率去噪器($x$-prediction)和噪声预测($\epsilon$-prediction))的通用公式从这些概率路径中进行采样;(iii)实际上,专注于使用不同调度程序定义的特定概率路径,相比以前的离散扩散和流模型,可以显着提高生成的困惑度;(iv)通过将离散流匹配模型扩展到17亿个参数,我们在HumanEval上达到了6.7% Pass@1和13.4% Pass@10,在1-shot MBPP编码基准测试上达到了6.7% Pass@1和20.6% Pass@10。我们的方法能够以非自回归的方式生成高质量的离散数据,显著缩小了自回归模型和离散流模型之间的差距。
Despite Flow Matching and diffusion models having emerged as powerful generative paradigms for continuous variables such as images and videos, their application to high-dimensional discrete data, such as language, is still limited. In this work, we present Discrete Flow Matching, a novel discrete flow paradigm designed specifically for generating discrete data. Discrete Flow Matching offers several key contributions: (i) it works with a general family of probability paths interpolating between source and target distributions; (ii) it allows for a generic formula for sampling from these probability paths using learned posteriors such as the probability denoiser ($x$-prediction) and noise-prediction ($\epsilon$-prediction); (iii) practically, focusing on specific probability paths defined with different schedulers considerably improves generative perplexity compared to previous discrete diffusion and flow models; and (iv) by scaling Discrete Flow Matching models up to 1.7B parameters, we reach 6.7% Pass@1 and 13.4% Pass@10 on HumanEval and 6.7% Pass@1 and 20.6% Pass@10 on 1-shot MBPP coding benchmarks. Our approach is capable of generating high-quality discrete data in a non-autoregressive fashion, significantly closing the gap between autoregressive models and discrete flow models.
许愿