Train agents using value functions, policy methods, temporal-difference learning, actor-critic algorithms, and environment simulation.