Diffusion Models
A denoising diffusion probabilistic model defines a fixed process that gradually corrupts data and learns a reverse process that reconstructs samples step by step.
Open full-size imageCheck the direction before reading the formulas: this picture places noisy xT on the left and clean x0 on the right. The solid arrows show learned denoising p; the dashed arrow q points back toward more noise. Training samples noisy states; generation follows the reverse transitions.
Forward Process
For a variance schedule and ,
Let . Any timestep can be sampled directly from clean data:
This closed form makes it unnecessary to simulate every earlier noising step during training.
Learned Reverse Process
Generation needs the reverse transition
One common parameterization predicts the noise added to . A widely used simplified objective is
This noise-prediction loss is connected to the variational objective, but exact weighting and parameterization matter; “predict noise with MSE” is not the complete definition of every diffusion model.
Sampling
- Draw from the chosen noise distribution.
- For , use the learned model and sampler to obtain .
- Return the final representation.
Conditioning supplies additional information—such as a class, text representation, or measurement—to the denoiser. Guidance can strengthen conditioning at the cost of changing the diversity, fidelity, or calibration trade-off.
Turn Noise Prediction into a Reverse Step
For the discrete Gaussian DDPM in the original paper, Algorithms 1–2, take and . Choose a schedule with close to zero so the terminal distribution is approximately standard Gaussian. Starting from is otherwise a mismatch, not a consequence of merely having many steps.
With an epsilon-predicting network, the reverse mean is
One fixed-variance choice is . Sample with fresh for ; at , return the mean without fresh noise. This specifies one sampler, not all diffusion solvers. The exact forward posterior conditioned on both and is Gaussian, but the reverse conditional without known generally is not; the learned Gaussian transition is a modeling approximation.
For a scalar check, let , and the sampled training noise be . Then . A prediction of has squared noise error and implies . Predicting the actual sampled noise would recover in this arithmetic example; from alone, that noise is not uniquely identifiable. MSE learns a conditional mean prediction, not an oracle inverse for each random draw.
Training samples data, a timestep (commonly uniformly from to ), and noise, then performs one denoiser evaluation and a gradient update. Generation has no clean target and repeatedly evaluates the denoiser with fixed weights. The noise target and denoiser output have the same shape as , and timestep conditioning tells the network which noise level it must handle. This note covers discrete Gaussian DDPM mechanics, not flow matching or every continuous-time and distilled sampler.
Design and Evaluation Boundaries
- The noise schedule, prediction target, model architecture, and sampler are separate design choices.
- Standard sampling is iterative and can require many model evaluations; faster samplers trade computation against approximation behavior.
- A low denoising objective does not by itself establish perceptual quality, diversity, likelihood quality, or usefulness.
- Conditional generation can reproduce biases or memorized structure from training data.
- The model's output space, preprocessing, and decoder can be as important as the denoiser.
- Current image, audio, or API products belong in Frontier; this note owns the durable probabilistic process.
The canonical starting point is Denoising Diffusion Probabilistic Models. Framework-specific U-Net and sampling implementations are intentionally left to maintained libraries and papers.
Diffusion Explainer follows preset text prompts through Stable Diffusion’s text encoder, iterative denoising, and image decoder. Step through the process and compare prompts to locate where conditioning enters. This is a latent-diffusion application of the ideas above, rather than a visualization of exactly the pixel-space DDPM equations.