Pretraining, Supervised Fine-Tuning, and Preference Optimization
Explain how training targets shape representations, instruction following, and preferred behavior without equating preference with truth.
Explain how training targets shape representations, instruction following, and preferred behavior without equating preference with truth.
Follow one parameter update, then connect minibatches, optimizer state, validation, and reproducible checkpoints.