Skip to main content

Diffusion Models

A denoising diffusion probabilistic model defines a fixed process that gradually corrupts data and learns a reverse process that reconstructs samples step by step.

The learned reverse process runs from noise xT on the left toward data x0 on the right; the dashed arrow denotes forward noising.Open full-size image

Check the direction before reading the formulas: this picture places noisy xT on the left and clean x0 on the right. The solid arrows show learned denoising p; the dashed arrow q points back toward more noise. Training samples noisy states; generation follows the reverse transitions.

Forward Process​

For a variance schedule βt\beta_t and αt=1−βt\alpha_t=1-\beta_t,

q(xt∣xt−1)=N ⁣(xt;αt xt−1,(1−αt)I).q(x_t\mid x_{t-1}) = \mathcal{N}\!\left( x_t;\sqrt{\alpha_t}\,x_{t-1}, (1-\alpha_t)I \right).

Let αˉt=∏s=1tαs\bar{\alpha}_t=\prod_{s=1}^{t}\alpha_s. Any timestep can be sampled directly from clean data:

xt=αˉt x0+1−αˉt ϵ,ϵ∼N(0,I).x_t = \sqrt{\bar{\alpha}_t}\,x_0 + \sqrt{1-\bar{\alpha}_t}\,\epsilon, \qquad \epsilon\sim\mathcal{N}(0,I).

This closed form makes it unnecessary to simulate every earlier noising step during training.

Learned Reverse Process​

Generation needs the reverse transition

pθ(xt−1∣xt)=N ⁣(xt−1;μθ(xt,t),Σθ(xt,t)).p_{\theta}(x_{t-1}\mid x_t) = \mathcal{N}\!\left( x_{t-1};\mu_{\theta}(x_t,t), \Sigma_{\theta}(x_t,t) \right).

One common parameterization predicts the noise added to x0x_0. A widely used simplified objective is

Ex0,t,ϵ[∥ϵ−ϵθ(αˉtx0+1−αˉtϵ,t)∥22].\mathbb{E}_{x_0,t,\epsilon} \left[ \left\| \epsilon-\epsilon_{\theta} \left( \sqrt{\bar{\alpha}_t}x_0 +\sqrt{1-\bar{\alpha}_t}\epsilon, t \right) \right\|_2^2 \right].

This noise-prediction loss is connected to the variational objective, but exact weighting and parameterization matter; “predict noise with MSE” is not the complete definition of every diffusion model.

Sampling​

  1. Draw xTx_T from the chosen noise distribution.
  2. For t=T,…,1t=T,\ldots,1, use the learned model and sampler to obtain xt−1x_{t-1}.
  3. Return the final x0x_0 representation.

Conditioning supplies additional information—such as a class, text representation, or measurement—to the denoiser. Guidance can strengthen conditioning at the cost of changing the diversity, fidelity, or calibration trade-off.

Turn Noise Prediction into a Reverse Step​

For the discrete Gaussian DDPM in the original paper, Algorithms 1–2, take 0<βt<10<\beta_t<1 and αˉ0=1\bar\alpha_0=1. Choose a schedule with αˉT\bar\alpha_T close to zero so the terminal distribution is approximately standard Gaussian. Starting from N(0,I)\mathcal N(0,I) is otherwise a mismatch, not a consequence of merely having many steps.

With an epsilon-predicting network, the reverse mean is

μθ(xt,t)=1αt(xt−βt1−αˉtϵθ(xt,t)).\mu_\theta(x_t,t)=\frac{1}{\sqrt{\alpha_t}} \left(x_t-\frac{\beta_t}{\sqrt{1-\bar\alpha_t}}\epsilon_\theta(x_t,t)\right).

One fixed-variance choice is β~t=βt(1−αˉt−1)/(1−αˉt)\tilde\beta_t=\beta_t(1-\bar\alpha_{t-1})/(1-\bar\alpha_t). Sample xt−1=μθ+β~t ξx_{t-1}=\mu_\theta+\sqrt{\tilde\beta_t}\,\xi with fresh ξ∼N(0,I)\xi\sim\mathcal N(0,I) for t>1t>1; at t=1t=1, return the mean without fresh noise. This specifies one sampler, not all diffusion solvers. The exact forward posterior conditioned on both xtx_t and x0x_0 is Gaussian, but the reverse conditional without known x0x_0 generally is not; the learned Gaussian transition is a modeling approximation.

For a scalar check, let αˉt=0.64\bar\alpha_t=0.64, x0=2x_0=2 and the sampled training noise be ϵ=1\epsilon=1. Then xt=0.8(2)+0.6(1)=2.2x_t=0.8(2)+0.6(1)=2.2. A prediction of 0.50.5 has squared noise error 0.250.25 and implies x^0=(2.2−0.6(0.5))/0.8=2.375\hat x_0=(2.2-0.6(0.5))/0.8=2.375. Predicting the actual sampled noise would recover 22 in this arithmetic example; from xtx_t alone, that noise is not uniquely identifiable. MSE learns a conditional mean prediction, not an oracle inverse for each random draw.

Training samples data, a timestep (commonly uniformly from 11 to TT), and noise, then performs one denoiser evaluation and a gradient update. Generation has no clean target and repeatedly evaluates the denoiser with fixed weights. The noise target and denoiser output have the same shape as xtx_t, and timestep conditioning tells the network which noise level it must handle. This note covers discrete Gaussian DDPM mechanics, not flow matching or every continuous-time and distilled sampler.

Design and Evaluation Boundaries​

  • The noise schedule, prediction target, model architecture, and sampler are separate design choices.
  • Standard sampling is iterative and can require many model evaluations; faster samplers trade computation against approximation behavior.
  • A low denoising objective does not by itself establish perceptual quality, diversity, likelihood quality, or usefulness.
  • Conditional generation can reproduce biases or memorized structure from training data.
  • The model's output space, preprocessing, and decoder can be as important as the denoiser.
  • Current image, audio, or API products belong in Frontier; this note owns the durable probabilistic process.

The canonical starting point is Denoising Diffusion Probabilistic Models. Framework-specific U-Net and sampling implementations are intentionally left to maintained libraries and papers.

Diffusion Explainer follows preset text prompts through Stable Diffusion’s text encoder, iterative denoising, and image decoder. Step through the process and compare prompts to locate where conditioning enters. This is a latent-diffusion application of the ideas above, rather than a visualization of exactly the pixel-space DDPM equations.

Explore connectionsOpen network