Skip to main content

Variational Autoencoders

A variational autoencoder (VAE) is a latent-variable generative model. It specifies a prior p(z)p(z) and a decoder likelihood pθ(x∣z)p_{\theta}(x\mid z):

pθ(x)=∫p(z) pθ(x∣z) dz.p_{\theta}(x) = \int p(z)\,p_{\theta}(x\mid z)\,dz.

Given an observation xx, the posterior pθ(z∣x)p_{\theta}(z\mid x) describes which latent values could explain it under the model. This distribution is usually difficult to compute directly, so an encoder qϕ(z∣x)q_{\phi}(z\mid x) approximates it. The same encoder, with shared parameters ϕ\phi, predicts an approximate posterior for each observation. This reuse across observations is called amortized inference.

Evidence Lower Bound​

Training maximizes the evidence lower bound (ELBO):

L(x;θ,ϕ)=Eqϕ(z∣x)[log⁡pθ(x∣z)]−KL ⁣(qϕ(z∣x) ∥ p(z)).\mathcal{L}(x;\theta,\phi) = \mathbb{E}_{q_{\phi}(z\mid x)} \left[\log p_{\theta}(x\mid z)\right] - \mathrm{KL}\!\left(q_{\phi}(z\mid x)\,\|\,p(z)\right).

It is a lower bound because

log⁡pθ(x)=L(x;θ,ϕ)+KL ⁣(qϕ(z∣x) ∥ pθ(z∣x)).\log p_{\theta}(x) = \mathcal{L}(x;\theta,\phi) + \mathrm{KL}\!\left( q_{\phi}(z\mid x) \,\|\, p_{\theta}(z\mid x) \right).

Within the ELBO, the expected decoder log-likelihood rewards explaining the observation. Its prior KL penalty keeps the approximate posterior near the prior, helping samples drawn from the prior reach regions the decoder has learned; this does not guarantee generation quality. Calling the expected log-likelihood merely “reconstruction loss” can hide an important modeling choice: Bernoulli, Gaussian, categorical, and other likelihoods imply different objectives and data assumptions.

Reparameterized Gradients​

For a diagonal Gaussian encoder,

qϕ(z∣x)=N ⁣(z;μϕ(x),diag⁡(σϕ2(x))),q_{\phi}(z\mid x) = \mathcal{N}\!\left(z;\mu_{\phi}(x), \operatorname{diag}(\sigma_{\phi}^2(x))\right),

sample through parameter-free noise:

ϵ∼N(0,I),z=μϕ(x)+σϕ(x)⊙ϵ.\epsilon\sim\mathcal{N}(0,I), \qquad z=\mu_{\phi}(x)+\sigma_{\phi}(x)\odot\epsilon.

This moves randomness outside the differentiable path through μϕ\mu_{\phi} and σϕ\sigma_{\phi}.

For a standard normal prior, the KL term has a closed form:

KL(q∥p)=−12∑j(1+log⁡σj2−μj2−σj2).\mathrm{KL}(q\|p) = -\frac{1}{2}\sum_j \left(1+\log\sigma_j^2-\mu_j^2-\sigma_j^2\right).

One Observation, One Training Loss​

Here KL(q∥p)=Eq[log⁡q−log⁡p]\mathrm{KL}(q\|p)=E_q[\log q-\log p] measures distributional mismatch; it is nonnegative, asymmetric, and zero only when the distributions agree almost everywhere. The ELBO identity assumes the relevant densities and expectations are well-defined. It follows by substituting Bayes' rule pθ(z∣x)=pθ(x∣z)p(z)/pθ(x)p_\theta(z\mid x)=p_\theta(x\mid z)p(z)/p_\theta(x) into that KL definition. The gap vanishes when the encoder equals the true posterior, not merely when reconstruction is good. This is the variational argument in Auto-Encoding Variational Bayes.

Take a one-dimensional encoder with μ=1\mu=1, σ=0.5\sigma=0.5 and a fixed noise draw ϵ=0.2\epsilon=0.2, giving z=1.1z=1.1. Suppose the decoder is p(x∣z)=N(z,1)p(x\mid z)=\mathcal N(z,1) and the observed value is x=2x=2. Its negative log-likelihood at this draw is (2−1.1)2/2+log⁡(2π)/2≈1.3239(2-1.1)^2/2+\log(2\pi)/2\approx1.3239. The analytic KL to N(0,1)\mathcal N(0,1) is (1+0.25−1−log⁡0.25)/2≈0.8181(1+0.25-1-\log0.25)/2\approx0.8181. This one-sample negative-ELBO estimate is therefore 2.14212.1421. It estimates the expectation; a particular Monte Carlo draw is not itself guaranteed to bound log⁡p(x)\log p(x).

For a batch of BB observations and latent width dzd_z, the encoder emits mean and log-variance arrays of shape [B, d_z]. Set σ=exp⁡(logvar/2)\sigma=\exp(\text{logvar}/2), sample noise of the same shape, sum reconstruction log-likelihood over observation dimensions and KL over latent dimensions, then average over the batch. Averaging only the reconstruction dimensions instead of summing silently changes its weight relative to KL. Generation uses a prior draw and the decoder without the encoder; returning the decoder mean is a deterministic summary, not a full likelihood sample.

Generation and Variants​

  • Generation: sample z∼p(z)z\sim p(z), then sample or decode x∼pθ(x∣z)x\sim p_{\theta}(x\mid z).
  • Conditional VAE: condition the encoder and decoder on additional information yy.
  • β\beta-VAE: weight the KL term by β\beta to alter the rate–distortion trade-off. A larger weight does not guarantee a uniquely disentangled or semantically meaningful representation.
Digits generated across a two-dimensional latent-space grid change shape smoothly between neighboring positions.Open full-size image

Each tile decodes one latent position. Follow a row or column to see digit shapes change. The paper obtains coordinates through Gaussian inverse-CDF mapping. This is a grid through a learned latent space, not a gallery of random independent samples; smooth changes do not prove disentangled concepts.

Failure Modes​

  • Posterior collapse: a strong decoder may ignore zz, leaving qϕ(z∣x)q_{\phi}(z\mid x) close to the prior.
  • Approximation gap: the chosen variational family may poorly represent the true posterior.
  • Likelihood mismatch: a convenient decoder distribution can produce undesirable sample or reconstruction behavior.
  • Latent interpretation: smooth interpolation does not prove that latent coordinates correspond to independent human concepts.
  • Anomaly detection: reconstruction or likelihood scores can fail on unfamiliar data and need task-specific validation.

The canonical source is Auto-Encoding Variational Bayes. Framework training code is intentionally omitted because the model assumptions and ELBO are the durable part.

Explore connectionsOpen network