From SFT to DPO to RLHF: a Deep Guide to LLM Alignment & Preference Optimization
Why this exists Pretraining teaches a model to continue text; alignment teaches it to behave the way users want (helpful, honest, harmless, on-style, and on-policy). We do that with post-training: supervised data, preferences, rules, and sometimes re...
Sep 22, 20259 min read29
