Recurrent Neural Networks
An RNN reuses parameters while carrying state through time:
The finite-dimensional state is a learned summary for the objective, not a lossless archive. That compression enables streaming and also loses long-range detail.
Read horizontally to follow the hidden state through time, then vertically to connect an input to its output. Repeated boxes represent repeated use of shared parameters, not separately trained networks. A later output can depend on earlier inputs through this state path.
Sequence Objectives
Autoregressive modeling factorizes
Teacher forcing supplies the true prefix in training; free generation uses the model's own earlier outputs, including their errors. Low teacher-forced loss therefore does not guarantee stable rollout. This difference is often called exposure bias. Its remedies depend on the task; longer training alone does not eliminate it.
Why Backpropagation Through Time Fails
The path from time to contains a Jacobian product:
Repeated contraction causes vanishing gradients; expansion causes exploding gradients. Sequence length, weight spectra, activation saturation, and the trajectories encountered in the data all matter. The problem cannot be explained simply by calling RNNs old.
- Gradient clipping limits large updates but cannot restore vanished gradients.
- Truncated BPTT reduces memory and computation while truncating long-range credit assignment.
- Suitable initialization, normalization, and gates can improve training, without guaranteeing unlimited memory.
For clarity, define . The ordered product is ; matrix factors cannot generally be reversed. The same recurrent parameters appear at every step, so their gradient sums contributions from all uses, including indirect paths through later states. This is the chain-rule mechanism behind BPTT's gradient difficulties.
For the smoothing recurrence below with , , and inputs , states are . The sensitivity to after three steps is . Detaching the state after step two preserves its numerical value for the next forward step but removes the gradient path to the earlier chunk. Detaching is not resetting. For padded sequences, masking only the loss is insufficient if a final state is consumed: skip padded state updates or select the last valid state.
LSTM Gating
A long short-term memory network (LSTM) maintains a cell state alongside the hidden state . The forget gate scales the previous cell state, the input gate scales the candidate update , and the output gate controls how the cell state contributes to . Here is the sigmoid function, means elementwise multiplication, and concatenates the two vectors. In these equations, denotes the output gate, not the output scores in the opening RNN equation. A common form is:
The additive cell path improves information and gradient flow. Gates are learned soft controls, not automatically interpretable memory switches; implementations vary in biases, projections, peepholes, and gate order.
Worked Example: Streaming Smoothing
Consider the scalar recurrence
This is an exponential moving average. Each update uses constant memory. The newest observation has weight , and an observation steps old has weight .
The example shows both sides of recurrence: an RNN can update online without storing the whole sequence, but early information decays according to a fixed rule, and one old value cannot be reread on demand. A learned RNN is more flexible, yet remains constrained by its finite state and training signal. Attention changes this trade-off by retaining multiple represented positions and accessing them by content.
State and Data Boundaries
- Causal RNNs use only current and past inputs, so they support streaming prediction.
- Bidirectional RNNs also use future context, so they do not suit online decisions where that context is unavailable.
- Mask padding within a batch; sequence lengths revealed by padding must not leak labels.
- Reset hidden state between independent entities unless carrying it across the boundary has explicit meaning.
- Randomly splitting time-series windows can put adjacent or overlapping segments in both training and test.
The final item is data leakage rather than an architectural performance result. Design splits by entity and time.
Modern State-Based Sequence Models
Classical RNNs are not the end of state-based modeling. Structured state-space models (SSMs) start from continuous- or discrete-time linear state equations and use structured parameters and parallel algorithms to combine long-sequence training with recurrent inference. S4 demonstrated the long-sequence capabilities of structured SSMs; Mamba made parts of the update input-dependent and introduced a hardware-aware scan; xLSTM revisited recurrent memory and gates.
These names describe distinct architectures:
- A classical RNN's nonlinear state transition differs from an SSM's structured state equations.
- Mamba's linear sequence-scaling claim describes an algorithmic path, not guaranteed wall-clock superiority over attention at every length, batch size, or device.
- xLSTM continues the gated recurrent approach; it does not establish that LSTMs universally beat Transformers.
- Paper benchmarks remain tied to data, scale, kernels, and training budgets, and need replication on the target task.
As of 2026-08-11, S4, Mamba, and xLSTM are important competing designs, not a new permanent default.
RNNs, SSMs, and Full Self-Attention
Neither “Transformers made RNNs useless” nor “linear-time SSMs must replace attention” is justified. The choice depends on latency, memory, sequence length, training scale, kernels, and task structure.
Minimal Experiment
- Overfit a short sequence to check the implementation.
- Test whether changing masked padding leaves outputs unchanged.
- Deliberately mix state boundaries and check that the tests catch the contamination.
- Report teacher-forced loss and free-running results separately.
- Split by time and entity, and compare with stateless, seasonal, or moving-average baselines.
- Record gradient norms, truncation length, hidden width, kernels, and failed seeds.
This note favors supervised discrete sequences. Continuous-time systems, all SSM variants, reinforcement learning, and online feedback require separate treatment.
In Distill’s Visualizing memorization in RNNs, select characters in the autocomplete examples and inspect their connections to earlier input. Compare GRU and LSTM examples to see how similar aggregate scores can hide different uses of context. The displayed connections are gradient-based diagnostics for those examples, not a universal ranking of recurrent architectures.