Formulate states, actions, rewards, and policies; implement value-based and policy-based methods; and evaluate agents without unsafe trial and error.