Activations and Gated Feed-Forward Networks
“ReLU is obsolete” is too broad; “learn only ReLU, sigmoid, and tanh” is outdated. Choices now differ by architecture, scale, precision, kernels, and hardware. Transformer FFNs also often replace a plain two-layer path with a gated one, which changes more than an elementwise function.
The subject here is nonlinearity and channel mixing within a position; attention exchanges information across positions. An FFN gate multiplies channel values, whereas MoE routing chooses expert networks for tokens; a model can use both. The training process explains how a loss updates their parameters.
Elementwise Functions
is the standard normal cumulative distribution function. GELU may use the exact erf expression or a tanh approximation. Different kernels can produce small differences in mixed precision. The Swish paper studies learnable ; framework SiLU commonly fixes . Papers often use the names nearly interchangeably, so record the implementation.
Here and is the standard normal cumulative distribution function. Exact GELU uses ; the common tanh expression is an approximation. The table's form assumes ; unrestricted PReLU is defined piecewise as for nonnegative inputs and otherwise.
For SiLU, . At zero, ReLU, GELU, and SiLU all output zero, but GELU and SiLU have derivative ; ReLU has no unique derivative there (implementations choose a convention). At , SiLU is approximately and its derivative is approximately . A SwiGLU gate is therefore not a probability constrained to : it may be negative or exceed one.
Open full-size imageFocus on negative inputs: ReLU returns zero, GELU dips below zero then approaches it for very negative inputs, and ELU with α = 1 approaches −1. The blue curve is the standard GELU. Curve shape explains different responses, but does not rank their performance on a task.
Smooth Does Not Mean Universally Better
Smooth functions can improve some optimization trajectories and large-scale training results, but their benefits interact with initialization, normalization, width, data, optimization, and fused kernels. ReLU still provides three practical counterexamples to claims of universal progress:
- In convolutional, lightweight, or quantized deployments, simple kernels and exact zeros may matter more.
- With little data or small models, error is often dominated by the data and regularization.
- A stable ReLU baseline can reveal whether a complex gate's apparent benefit merely comes from added parameters.
A higher average score in a paper does not make a function the default for every task.
GLUs and Gated FFNs
A plain Transformer FFN is
A gated linear unit (GLU) multiplies two projections. A simplified form is:
Common Transformer variants include
ReGLU substitutes ReLU in the gate branch. A gated layer has three major projections (), whereas a plain FFN has only input and output projections. Keeping the same intermediate width increases both parameters and computation. Equal-parameter comparisons therefore usually reduce the gated layer’s intermediate width. The ratio depends on whether biases and embeddings are counted and on alignment constraints; no fixed ratio is universal.
For a bias-free FFN mapping width back to through width , the two matrices contain weights. A gated FFN of intermediate width contains weights, so equality gives . For , choose : both have 144 weights. The gate and value projections each map 6 features to 8, their elementwise product is still length 8, and the output projection maps back to 6. This is the local parameter-count argument used in GLU Variants Improve Transformer, not a guarantee of equal measured latency.
Dated Architecture Examples
- The original BERT architecture uses GELU.
- PaLM reports using SwiGLU.
- The Llama 3 architecture report also uses SwiGLU, alongside RMSNorm, RoPE, and GQA as parts of a complete block.
- ConvNeXt shows that modern ConvNets can use GELU too; smooth activations are not exclusive to Transformers.
These are examples documented in papers, not proof that GELU or SwiGLU wins at every scale, modality, or hardware setting. When several components change together, the overall benefit cannot be attributed entirely to activation.
Output Activations Are a Separate Decision
- Binary classification commonly trains one logit and applies sigmoid when interpreting it as a probability.
- Exclusive multiclass classification uses logits with Softmax.
- Multilabel classification uses independent sigmoids.
- Regression may use a linear, positive-valued, or bounded transform, depending on the target's support.
For numerical stability, training losses generally consume logits directly rather than manually applying sigmoid/Softmax and then taking logarithms.
Failure Modes
Controlled Selection
- Fix the data split, initialization family, optimizer, number of steps, and stopping rule first.
- When comparing plain ReLU/GELU/SiLU FFNs, keep widths and parameter counts interpretable and explicit.
- When comparing GEGLU/SwiGLU, report intermediate width, total parameters, FLOPs, and peak memory.
- Repeat multiple seeds and report means, dispersion, and failed runs.
- Measure end-to-end throughput at the target dtype, compiler, and hardware.
- Set quality, latency, memory, and stability thresholds before seeing results, rather than choosing metrics afterward.
Update Rules and Bias
This note focuses on dense networks, Transformers, and GPU training. Spiking models, periodic activations, implicit neural representations, learnable splines, symbolic or logic gates, and specialized accelerators may require entirely different selection criteria. Ask whether a new function changes representation, optimization, budget, or kernel before adding it to an activation zoo.