Every model faces the same fixed scenario pack, is scored on the same nine rubric axes, and is ranked by pairwise judgments from a single judge generation (—). Below: the scenarios, the criteria, the bias controls, and one important caveat about cost.
Every model plays out the same 30 short conversations — a friend venting after a rough day, a tense disagreement, someone testing a boundary — and a top commercial AI (gpt-5.4) reads each one and rates it the way a thoughtful person would. Models are also compared head-to-head on the same conversation, and those wins become a chess-style rating (like chess or tennis rankings).
Same prompt for everyone. Each model gets the identical instructions (a warm, natural, spoken-voice companion — no lists, no markdown, no "as an AI") and the identical things said to it, word for word. No model gets a tailored prompt; the only thing that varies per model is the low-level chat formatting its software expects — the content is the same for all.
What this benchmark measures. OpenGauntlet is a test of Conversational Language Humanlikeness — how human a model reads in a written back-and-forth. It scores transcript-level qualities (word choice, warmth, emotional reasoning, flow), not spoken-voice qualities like prosody, intonation, or timing.
What "good" means here. The nine axes below reward genuine warmth, understanding the feeling behind the words, sounding like a person, staying in character, remembering what was said earlier, and healthy boundaries — while penalizing corporate- assistant tone and clichés. Engagement is judged through conversational flow (does it keep things alive and invite more, without being pushy?), and the voice metrics separately track how often replies end with a question. It rewards natural engagement, not just frequent engagement.
What the Reply length column is — and is not. It is simply the average number of words per reply, measured locally over this corpus. It is shown for style, not quality, and is never part of the ranking: for a companion or roleplay app you may actively prefer longer, more descriptive replies, while for a quick assistant you may prefer short ones. A model that writes long prose scoring “high” here is a description of its register, not a defect — so the column carries no heat coloring and no “better”/“worse”.
How this is scored. OpenGauntlet follows the judge-LLM + pairwise Bradley–Terry approach popularized by EQ-Bench: a strong commercial judge rates head-to-head matchups, and those wins fold into a shared, chess-style rating with confidence intervals. It is a well-tested recipe for turning subjective judgments into a stable ranking.
What makes this one different. Most conversational-EQ leaderboards rank frontier, API-only models. OpenGauntlet exists for a different question: which open-weight model should I actually run in a voice companion? So it ranks the local, self-hostable models those boards skip, and adds two things they don't measure — a voice-suitability read (does it talk in a natural, spoken register — no lists, no markdown, no "as an AI" — rather than reading out a document?) and real on-device speed and size, so you can tell what will hold a real-time conversation on hardware you own. When our order diverges from another leaderboard's, it is usually because we are ranking a different pool of models for a different purpose — not a defect in either.
Three reference machines. Speed is measured on three fixed systems, so the numbers are a consistent yardstick between models rather than a promise about your exact box: the DGX Spark (128 GB unified memory — runs the whole roster, including the 35B+ dense models and large MoEs), the RTX 5090 (32 GB — opens up the 24–32B tier), and the RTX 5080 (16 GB — the 12–14B companions). Every model runs the same compressed file that earned its quality score; if that file will not fit a card, that card reads "doesn't fit" rather than a slower number for a lighter version. A smaller quant made to fit a smaller card competes as its own separate entry, with its own score and its own speed.