Form hypotheses, choose outcome and guardrail metrics, estimate sample needs, and run controlled experiments. Interpret significance, practical impact, novelty effects, and failed tests without cherry-picking.