Safety, alignment, and red teaming: Generative AI course | Zoonk
54. Safety, alignment, and red teaming
Reduce harmful, deceptive, biased, or dangerous outputs through policy design, safety training, classifiers, guardrails, red teaming, and staged release. Measure both over-refusal and under-refusal, with escalation paths for high-risk use.