//opengauntlet
// head to head

Put two models side by side

Pick any two models to line up their headline metrics, per-criterion scores, the judge's head-to-head pairwise verdict, and — where they answered the same scenario — their actual transcripts, turn for turn.

Loading models…