Pick any two models to line up their headline metrics, per-criterion scores, the judge's head-to-head pairwise verdict, and — where they answered the same scenario — their actual transcripts, turn for turn.