Skip to main content

Neural and generative models

Architecture determines how information is represented and combined. A training objective says what behavior is rewarded; a decoding procedure says how outputs are obtained. These are connected choices, not synonyms.

Follow MLPs and activations into convolutional and recurrent alternatives. For text, read tokenization, attention, Transformers and encoder/decoder families before generation. VAE and diffusion explain other generative routes; multimodal models and MoE extend representations and computation.

Reading Order​

StepArticleWhat it explains
1Multilayer PerceptronUnderstand MLP capacity, backpropagation, optimization failures, and inductive bias through tensor shapes and a worked XOR construction.
2Activations and Gated Feed-Forward NetworksDistinguish ReLU, GELU, SiLU/Swish, and GLU variants, then select nonlinearities with controlled modern-network experiments.
3Convolutional Neural NetworksConvolutional inductive bias, channels, padding, stride, pooling, and the distinction between translation equivariance and invariance.
4Recurrent Neural NetworksUnderstand recurrent state compression, backpropagation through time, LSTM gating, and the boundary with modern state-space sequence models.
5Tokenization and Text Representations Inside a ModelFollow text through token IDs, embeddings, and contextual states, distinguishing tokenization, position, padding, and retrieval vectors.
6Attention MechanismUnderstand content addressing, Q/K/V, masks, multi-head variants, efficient implementations, and interpretation limits through a numerical example.
7Transformers: Original Architecture and Modern BlocksDistinguish the 2017 encoder–decoder from common Pre-Norm, RMSNorm, RoPE, GQA, and SwiGLU blocks, including training and inference costs.
8Encoder, Decoder, and Encoder–Decoder: How Information FlowsDistinguish model families by visible context, training objectives, and output heads, from BERT classifiers to generators and sequence converters.
9Autoregressive Generation and DecodingFollow next-token distributions to complete answers, including temperature, top-k, top-p, search, stopping, and structured output.
10Variational AutoencodersLatent-variable generative models trained with amortized variational inference, the ELBO, and reparameterized gradients.
11Diffusion ModelsForward noising, learned reverse transitions, noise-prediction training, iterative sampling, conditioning, and computational trade-offs.
12Multimodal Models: Alignment, Fusion, and Visual QuestionsUse image–text retrieval and visual question answering to distinguish objectives, patches, modality connectors, information loss, and evaluation.
13Mixture of Experts: Sparse Computation, Routing, and Deployment CostFollow a token through MoE routing to distinguish total parameters, active parameters, memory, speed, load balancing, and communication.

Use What You Read​

Track tensor shapes and information access through one example. For a new architecture, ask what can attend to what, which parameters are shared, and what the forward pass actually returns. Training and inference costs have their own companion topics.

Return to the AI reading paths to choose a neighboring topic.

Explore connectionsOpen network