Operate services with service-level objectives, error budgets, runbooks, on-call practices, incident command, and blameless reviews. Turn production failures into architectural improvements.