Identify failure modes such as reward hacking, deceptive outputs, uncontrolled actions, and loss of human control, then apply capability evaluations and layered safeguards.