Activations and Gated Feed-Forward Networks
Distinguish ReLU, GELU, SiLU/Swish, and GLU variants, then select nonlinearities with controlled modern-network experiments.
Distinguish ReLU, GELU, SiLU/Swish, and GLU variants, then select nonlinearities with controlled modern-network experiments.
Understand content addressing, Q/K/V, masks, multi-head variants, efficient implementations, and interpretation limits through a numerical example.
Follow next-token distributions to complete answers, including temperature, top-k, top-p, search, stopping, and structured output.
Convolutional inductive bias, channels, padding, stride, pooling, and the distinction between translation equivariance and invariance.
Distinguish model families by visible context, training objectives, and output heads, from BERT classifiers to generators and sequence converters.
Follow a token through MoE routing to distinguish total parameters, active parameters, memory, speed, load balancing, and communication.
Understand MLP capacity, backpropagation, optimization failures, and inductive bias through tensor shapes and a worked XOR construction.
Use image–text retrieval and visual question answering to distinguish objectives, patches, modality connectors, information loss, and evaluation.
How do architectures represent inputs and produce outputs?
Understand recurrent state compression, backpropagation through time, LSTM gating, and the boundary with modern state-space sequence models.
Follow text through token IDs, embeddings, and contextual states, distinguishing tokenization, position, padding, and retrieval vectors.
Distinguish the 2017 encoder–decoder from common Pre-Norm, RMSNorm, RoPE, GQA, and SwiGLU blocks, including training and inference costs.