// one canonical page per measured configuration
Conversational AI model benchmark results
Every model and quantized configuration tested by OpenGauntlet, with its own rubric scores, pairwise standing, transcripts, and measured hardware performance.
All measured model configurations
- glistening-gem-31b (Q4_K_M) — rank #1, humanlikeness 77.9/100
- Gemma 4 31B StyleTune (Q4_K_M, no thinking) — rank #2, humanlikeness 77.5/100
- glistening-gem-31b — rank #3, humanlikeness 78.2/100
- Gemma 4 31B StyleTune (Q4_K_M) — rank #4, humanlikeness 78.2/100
- Gemma 4 31B LilaRest NVFP4 Turbo — rank #5, humanlikeness 76.6/100
- glistening-gem-31b v2.0 (Q4_K_M) — rank #6, humanlikeness 76.2/100
- Gemma 4 31B Ortenzya Creative Wordsmith Q8 — rank #7, humanlikeness 75.7/100
- skyfall-31b — rank #8, humanlikeness 72.0/100
- glistening-gem-31b (Q4_K_M, no thinking) — rank #9, humanlikeness 76.7/100
- g4-meromero-v2-31b (Q4_K_M) — rank #10, humanlikeness 76.2/100
- Gemma 4 26B-A4B StyleTune V2 (Q4_K_M) — rank #11, humanlikeness 71.0/100
- gemma4-26b-a4b-it — rank #12, humanlikeness 70.1/100
- Gemma 4 31B Dark Thoughts V2 (Q4_K_M) — rank #13, humanlikeness 74.8/100
- Gemma 4 31B it scotoma-2 (Q4_K_M) — rank #14, humanlikeness 76.9/100
- Gemma 4 26B-A4B Unsloth NVFP4 — rank #15, humanlikeness 67.4/100
- Qwen 3.6 27B Architect Polaris Fable NVFP4 — rank #16, humanlikeness 67.1/100
- qwen36-35b-styletune — rank #17, humanlikeness 70.5/100
- Versipellis 31B (Q4_K_M, no thinking) — rank #18, humanlikeness 76.3/100
- gemma4-12b-blorbo — rank #19, humanlikeness 69.6/100
- heretic-neocode-27b — rank #20, humanlikeness 63.8/100
- fable-fusion-27b — rank #21, humanlikeness 68.8/100
- g4-meromero-31b — rank #22, humanlikeness 71.7/100
- skyfall-31b (Q4_K_M) — rank #23, humanlikeness 69.1/100
- Gemma 4 31B Dark Thoughts V2 (Q4_K_M, no thinking) — rank #24, humanlikeness 74.3/100
- hivemind-32b-preview — rank #25, humanlikeness 63.6/100
- cydonia-24b-v4.3 — rank #26, humanlikeness 64.3/100
- Gemma 4 26B-A4B Huihui QAT Abliterated MTP NVFP4 — rank #27, humanlikeness 62.7/100
- Gemma 4 26B-A4B AEON Uncensored NVFP4 — rank #28, humanlikeness 64.3/100
- Gemma 4 26B-A4B Heretic QAT NVFP4 GGUF — rank #29, humanlikeness 64.6/100
- qwen3-14b — rank #30, humanlikeness 58.6/100
- qwen3.5-27b — rank #31, humanlikeness 53.3/100
- mythos-distilled-27b — rank #32, humanlikeness 56.7/100
- rocinante-xl-16b — rank #33, humanlikeness 65.9/100
- ornith-35b-nvfp4 — rank #34, humanlikeness 59.9/100
- anubis-mini-8b — rank #35, humanlikeness 50.2/100
- bluestar-27b — rank #36, humanlikeness 43.1/100
- rocinante-x-12b — rank #37, humanlikeness 50.3/100
- Magistry 24B v1.1 (Q4_K_M) — rank #38, humanlikeness 47.7/100
- NVIDIA Nemotron 3 Super 120B-A12B NVFP4 — rank #39, humanlikeness 38.7/100
- pantheon-proto-rp-1.8-30b-a3b — rank #40, humanlikeness 41.1/100
- designant — rank #41, humanlikeness 28.1/100
Results use the same scenario pack and judging protocol. Read the methodology or compare two models head to head.