Evaluate generated outputs with task-specific test sets, model-based graders, human review, pairwise comparisons, red-team cases, and statistical reporting.