Clustering and Dimensionality Reduction
Use worked k-means and PCA examples to distinguish discovering groups from compressing representations, and explain the limits of distance, scaling, and visualization.
Use worked k-means and PCA examples to distinguish discovering groups from compressing representations, and explain the limits of distance, scaling, and visualization.
A foundation map for convex sets, convex functions, duality, and optimization problems with global guarantees.
Derive a tree split and a boosting update, then compare their inductive biases, validation needs, and limits.
A foundation map for measuring uncertainty, information, compression limits, and distributional difference.
A foundation map for vectors, linear transformations, matrix factorization, and data-oriented applications.
Gaussian generative classifiers whose shared or class-specific covariance assumptions produce linear or quadratic decision boundaries.
Build a linear predictor, calculate residuals, and follow gradient updates; closed-form estimation and statistical inference are covered separately.
Why logarithmic loss measures probabilistic classification error and how calculus connects it to likelihood optimization.
What changes during training, and why does it generalize?
A compact vocabulary for models, losses, empirical risk, likelihood, regularization, optimization, and generalization.
How do architectures represent inputs and produce outputs?
Explain how training targets shape representations, instruction following, and preferred behavior without equating preference with truth.
Distinguish prediction accuracy, reliable probabilities, and worthwhile actions, and use independent data to set calibration and abstention rules.
Use a two-state example to distinguish immediate rewards, long-term value, exploration, and learned agent policies.
Multiclass linear classification with logits, softmax probabilities, cross-entropy, and clear boundaries around multilabel tasks.
Follow one parameter update, then connect minibatches, optimizer state, validation, and resumable checkpoints.
Choose between a frozen encoder, full fine-tuning, low-rank adaptation, and a smaller student for a defined task.