Colored Noise Diffusion Sampling

GenAI Diffusion ML DM
扩散模型在图像合成任务中达到了当前最优性能,其生成轨迹本质上呈现出一种谱偏差(spectral bias):即早期优先重建低频的全局结构,后期才逐步恢复高频的精细细节。传统的随机微分方程(SDE)求解器未能考虑这一动态演化特性,而是在整个生成过程中简单粗暴地注入均匀的白噪声,从而错误地消耗了本就有限的能量预算。本文构建了一个全新的数学框架,将SDE推理过程重新诠释为一种目标明确、按频率解耦的能量传递机制。基于该框架,我们提出了“彩色噪声采样”(Colored Noise Sampling, CNS)——一种新颖的、无需额外训练的随机求解器。CNS不再注入均匀白噪声,而是采用一种动态的、随时间步长与频率变化而自适应调整的噪声调度策略,从而更高效地将注入的能量定向分配至当前尚未被结构化重建的频率成分上。通过主动利用扩散模型固有的谱偏差特性,CNS系统性地引导生成分布向真实数据流形收敛。大量实验表明,作为一种严格即插即用、仅在推理阶段替换采样器的方法,CNS在多种主流架构(SiT、JiT、FLUX)上均显著优于标准的常微分方程(ODE)和随机微分方程(SDE)基线方法。在ImageNet-256数据集上的无分类器引导(unguided)采样实验中,CNS实现了显著的FID指标下降:SiT-XL/2模型的FID从8.26降至6.27;JiT-B/16模型从32.39降至26.69;JiT-H/16模型则从11.88降至8.31。此外,在采用分类器自由引导(Classifier-Free Guidance)时,CNS亦能保持稳定且一致的相对FID提升。项目主页见:https://hadardavidson.github.io/CNS/。
Diffusion models achieve state-of-the-art image synthesis, with their generative trajectories fundamentally exhibiting a spectral bias, resolving low-frequency global structures early and high-frequency fine details later. Conventional stochastic differential equation (SDE) solvers fail to account for this dynamic, naively injecting uniform white noise throughout the entire process and misusing the finite energy budget. In this work, we establish a mathematical framework that reconsiders SDE inference as a targeted, frequency-decoupled energy transfer. Leveraging this framework, we introduce Colored Noise Sampling (CNS), a novel, training-free stochastic solver. Rather than injecting uniform white noise, CNS utilizes a dynamic, timestep- and frequency-dependent schedule that more efficiently allocates injected energy toward structurally unresolved frequency bands. By actively exploiting the model's inherent spectral bias, CNS systematically steers the generated distribution toward the true data manifold. Extensive experiments demonstrate that CNS significantly outperforms standard ODE and SDE baselines as a strictly plug-and-play, inference-time sampler substitution across diverse architectures (SiT, JiT, FLUX). Compared to standard sampling on ImageNet-256, CNS achieves substantial unguided FID reductions, improving from 8.26 to 6.27 on SiT-XL/2, 32.39 to 26.69 on JiT-B/16, and 11.88 to 8.31 on JiT-H/16, while yielding consistent relative FID improvements with Classifier-Free Guidance. Project page is available at https://hadardavidson.github.io/CNS/.
许愿