//opengauntlet
// measured here, in text — free, open models you run yourself

Then, which one would you rather talk to?

Free models you can run on your own computer, ranked on one question: which of them is worth talking to? Same scenarios, same judge, same instructions for every model — and the whole ladder is below, most human at the top.

How human each one sounds

Every model we have run, on one scale. Higher is more human — the number is a normalised Elo from head-to-head judging, not a percentage.
Loading…

Graded by gpt-5.4, using the same instructions for every model.

//

The full table

Every column we record, for every model. Click a heading to sort.

Speed is in words per second, listed for every machine we measured a model on — a DGX Spark, an RTX 5090 and an RTX 5080. Most people read about 4 words a second, so anything above that is already arriving faster than you can read it. Machines a model is too big to fit on are left out.

Why some rows are the same model twice
Each row is one way of running a model. The same model compressed differently, or run with different software, is a separate row because it behaves differently. Each row links to its full record.
less human more human — color ranks each model within its own column; don't compare colors across different columns.
Loading…