Preference optimization and RLHF: Generative AI course | Zoonk
32. Preference optimization and RLHF
Align model behavior with human or model-generated preferences using reward models, RLHF, PPO, DPO, and related objectives. Address preference-data quality, reward hacking, sycophancy, and alignment tradeoffs.