Activations and Gated Feed-Forward Networks
Distinguish ReLU, GELU, SiLU/Swish, and GLU variants, then select nonlinearities with controlled modern-network experiments.
Distinguish ReLU, GELU, SiLU/Swish, and GLU variants, then select nonlinearities with controlled modern-network experiments.
Understand content addressing, Q/K/V, masks, multi-head variants, efficient implementations, and interpretation limits through a numerical example.
Convolutional inductive bias, channels, padding, stride, pooling, and the distinction between translation equivariance and invariance.
A path through MLPs, modern activations, RNNs, attention, and Transformers that also treats leakage and distribution shift as part of model evaluation.
A compact map of latent-variable and diffusion approaches to learning data distributions.
Understand MLP capacity, backpropagation, optimization failures, and inductive bias through tensor shapes and a worked XOR construction.
Understand recurrent state compression, backpropagation through time, LSTM gating, and the boundary with modern state-space sequence models.
Distinguish the 2017 encoder–decoder from common Pre-Norm, RMSNorm, RoPE, GQA, and SwiGLU blocks, including training and inference costs.