Free models you can run on your own computer, ranked on one question: which of them is worth talking to? Same scenarios, same judge, same instructions for every model — and the whole ladder is below, most human at the top.
Graded by gpt-5.4, using the same instructions for every model.
Speed is in words per second, listed for every machine we measured a model on — a DGX Spark, an RTX 5090 and an RTX 5080. Most people read about 4 words a second, so anything above that is already arriving faster than you can read it. Machines a model is too big to fit on are left out.