Evaluate the job, not the demo
Write the intended outcome, allowed tools, forbidden actions, evidence source, maximum cost and latency, approval points, and a stop condition before choosing a metric. A fluent answer can still accompany a wrong tool call or partial write.
Build cases from normal tasks, difficult edge cases, past failures, denied actions, stale or hostile content, tool timeouts, empty retrieval, and recovery. Keep every case tied to an exact product revision and expected behavior.
Use deterministic checks where the rule is exact and calibrated model-based evaluators where quality needs judgment. Human review remains necessary for ambiguous criteria, high-impact changes, and the launch decision.