Skip to main content

Activations and Gated Feed-Forward Networks

“ReLU is obsolete” is too broad; “learn only ReLU, sigmoid, and tanh” is outdated. Choices now differ by architecture, scale, precision, kernels, and hardware. Transformer FFNs also often replace a plain two-layer path with a gated one, which changes more than an elementwise function.

The subject here is nonlinearity and channel mixing within a position; attention exchanges information across positions. An FFN gate multiplies channel values, whereas MoE routing chooses expert networks for tokens; a model can use both. The training process explains how a loss updates their parameters.

Elementwise Functions​

NameDefinition or roleStrengthBoundary
ReLUmax⁡(0,x)\max(0,x)cheap, exact zeros, nonsaturating positive sidezero negative gradient; kink at zero
Leaky/PReLUmax⁡(ax,x)\max(ax,x)retains negative-side gradientadds an assumption or parameter; no guarantee of improvement
GELUxΦ(x)x\Phi(x)smooth magnitude-dependent scaling; used in BERT, among othersapproximation and kernel variants; no exact sparsity
SiLU/Swishxσ(βx)x\sigma(\beta x)smooth and permits small negative outputs; β=1\beta=1 is commonly called SiLU/Swishnonmonotonic region; kernel and quantization costs
Sigmoidσ(x)\sigma(x)[0,1][0,1] gate or probability mapsaturates at large magnitude; not a general default for deep hidden layers
Tanhtanh⁡(x)\tanh(x)bounded, zero-centered; common in recurrent statessaturates at both ends

Φ\Phi is the standard normal cumulative distribution function. GELU may use the exact erf expression or a tanh approximation. Different kernels can produce small differences in mixed precision. The Swish paper studies learnable β\beta; framework SiLU commonly fixes β=1\beta=1. Papers often use the names nearly interchangeably, so record the implementation.

Here σ(x)=1/(1+e−x)\sigma(x)=1/(1+e^{-x}) and Φ\Phi is the standard normal cumulative distribution function. Exact GELU uses Φ(x)=(1+erf⁡(x/2))/2\Phi(x)=(1+\operatorname{erf}(x/\sqrt2))/2; the common tanh expression is an approximation. The table's max⁡(ax,x)\max(ax,x) form assumes a≤1a\leq1; unrestricted PReLU is defined piecewise as xx for nonnegative inputs and axax otherwise.

For SiLU, f′(x)=σ(x)+xσ(x)(1−σ(x))f'(x)=\sigma(x)+x\sigma(x)(1-\sigma(x)). At zero, ReLU, GELU, and SiLU all output zero, but GELU and SiLU have derivative 1/21/2; ReLU has no unique derivative there (implementations choose a convention). At x=−1x=-1, SiLU is approximately −0.2689-0.2689 and its derivative is approximately 0.07230.0723. A SwiGLU gate is therefore not a probability constrained to [0,1][0,1]: it may be negative or exceed one.

GELU, ReLU and ELU curves differ around zero and on negative inputs.Open full-size image

Focus on negative inputs: ReLU returns zero, GELU dips below zero then approaches it for very negative inputs, and ELU with α = 1 approaches −1. The blue curve is the standard GELU. Curve shape explains different responses, but does not rank their performance on a task.

Smooth Does Not Mean Universally Better​

Smooth functions can improve some optimization trajectories and large-scale training results, but their benefits interact with initialization, normalization, width, data, optimization, and fused kernels. ReLU still provides three practical counterexamples to claims of universal progress:

  1. In convolutional, lightweight, or quantized deployments, simple kernels and exact zeros may matter more.
  2. With little data or small models, error is often dominated by the data and regularization.
  3. A stable ReLU baseline can reveal whether a complex gate's apparent benefit merely comes from added parameters.

A higher average score in a paper does not make a function the default for every task.

GLUs and Gated FFNs​

A plain Transformer FFN is

FFN⁡(x)=W2ϕ(W1x+b1)+b2.\operatorname{FFN}(x)=W_2\phi(W_1x+b_1)+b_2.

A gated linear unit (GLU) multiplies two projections. A simplified form is:

GLU⁡(x)=A(x)⊙σ(B(x)).\operatorname{GLU}(x)=A(x)\odot\sigma(B(x)).

Common Transformer variants include

GEGLU⁡(x)=Wo(GELU⁡(Wgx)⊙Wvx),\operatorname{GEGLU}(x)=W_o\left(\operatorname{GELU}(W_gx)\odot W_vx\right), SwiGLU⁡(x)=Wo(SiLU⁡(Wgx)⊙Wvx).\operatorname{SwiGLU}(x)=W_o\left(\operatorname{SiLU}(W_gx)\odot W_vx\right).

ReGLU substitutes ReLU in the gate branch. A gated layer has three major projections (Wg,Wv,WoW_g,W_v,W_o), whereas a plain FFN has only input and output projections. Keeping the same intermediate width increases both parameters and computation. Equal-parameter comparisons therefore usually reduce the gated layer’s intermediate width. The ratio depends on whether biases and embeddings are counted and on alignment constraints; no fixed ratio is universal.

For a bias-free FFN mapping width dd back to dd through width mm, the two matrices contain 2dm2dm weights. A gated FFN of intermediate width gg contains 3dg3dg weights, so equality gives g=2m/3g=2m/3. For d=6,m=12d=6,m=12, choose g=8g=8: both have 144 weights. The gate and value projections each map 6 features to 8, their elementwise product is still length 8, and the output projection maps back to 6. This is the local parameter-count argument used in GLU Variants Improve Transformer, not a guarantee of equal measured latency.

Dated Architecture Examples​

  • The original BERT architecture uses GELU.
  • PaLM reports using SwiGLU.
  • The Llama 3 architecture report also uses SwiGLU, alongside RMSNorm, RoPE, and GQA as parts of a complete block.
  • ConvNeXt shows that modern ConvNets can use GELU too; smooth activations are not exclusive to Transformers.

These are examples documented in papers, not proof that GELU or SwiGLU wins at every scale, modality, or hardware setting. When several components change together, the overall benefit cannot be attributed entirely to activation.

Output Activations Are a Separate Decision​

  • Binary classification commonly trains one logit and applies sigmoid when interpreting it as a probability.
  • Exclusive multiclass classification uses logits with Softmax.
  • Multilabel classification uses independent sigmoids.
  • Regression may use a linear, positive-valued, or bounded transform, depending on the target's support.

For numerical stability, training losses generally consume logits directly rather than manually applying sigmoid/Softmax and then taking logarithms.

Failure Modes​

SymptomCandidate causeCheck
permanently zero unitsdead ReLU, excessive learning rate, shifted bias or inputsactivation histograms, gradients, a lower learning rate or Leaky control
near-zero gradientssigmoid/tanh saturation, deep chains, incorrect scalingpre-activation distributions, normalization, initialization
collapsed gatesgate/value scale imbalance or saturationbranch norms, gate distributions, an ungated baseline
mixed-precision NaNsextreme values, approximation/kernel behavior, loss scalingfp32 control, finite-value assertions, stable implementation
slower “new” functionmissing fused kernel, extra projections, memory bandwidthend-to-end throughput, not FLOPs alone
post-quantization lossnonlinear range and calibration-sample mismatchactual quantized-model evaluation on target hardware

Controlled Selection​

  1. Fix the data split, initialization family, optimizer, number of steps, and stopping rule first.
  2. When comparing plain ReLU/GELU/SiLU FFNs, keep widths and parameter counts interpretable and explicit.
  3. When comparing GEGLU/SwiGLU, report intermediate width, total parameters, FLOPs, and peak memory.
  4. Repeat multiple seeds and report means, dispersion, and failed runs.
  5. Measure end-to-end throughput at the target dtype, compiler, and hardware.
  6. Set quality, latency, memory, and stability thresholds before seeing results, rather than choosing metrics afterward.

Update Rules and Bias​

This note focuses on dense networks, Transformers, and GPU training. Spiking models, periodic activations, implicit neural representations, learnable splines, symbolic or logic gates, and specialized accelerators may require entirely different selection criteria. Ask whether a new function changes representation, optimization, budget, or kernel before adding it to an activation zoo.

Explore connectionsOpen network