Learning track
Post-training & Alignment
Turn pretrained predictors into useful assistants through instruction tuning, preference learning, reinforcement learning, and distillation.
advanced · 11 available lessons
Post-training & Alignment
Turn pretrained predictors into useful assistants through instruction tuning, preference learning, reinforcement learning, and distillation.
- Why base models aren't assistants
- Supervised fine-tuning & instruction data
- LoRA, QLoRA & PEFT
- Reward models & human preference data
- RLHF with PPO
- DPO: skipping the reward model
- ORPO, KTO, SimPO & the alignment zoo
- GRPO & RL on verifiable rewards
- Reasoning models: test-time compute & long CoT
- Constitutional AI & RLAIF
- Distillation: making small models punch up