//opengauntlet
// surveyed, not benchmarked — 31 systems

LLM inference engines and runtimes compared

Every row on the leaderboard already names a serving backend, because the same weights produce different results depending on what runs them — different quantizations supported, different hardware backends, different throughput under load. Here is every LLM inference engine we could verify against a primary source, on the axes that decide a project — including a real benchmark OpenGauntlet ran on its own hardware, not just vendor claims.

Surveyed, not benchmarked. Nothing on this page was measured by OpenGauntlet's judge pipeline.

How to read this section

None of this is scored by the judge pipeline. OpenGauntlet measures Conversational Language Humanlikeness in text. This section is a sourced survey of inference engines, not a trial. Where a claim could not be verified against a primary source, it says so rather than smoothing it over.

Two rows are the exception. vLLM and SGLang cite a real benchmark OpenGauntlet ran on its own hardware — not a vendor's claim relayed from elsewhere. That evidence lives in their own row, clearly attributed; it does not change how this page as a whole is classified.

//

Every engine, compared

Filter by kind or by which quant formats it loads, pick the columns you care about, search any field, click a heading to sort. Every row carries the one thing a buyer would otherwise find out too late.
Loading…
commercial-safe weights permissive, with a catch not usable commercially proprietary service discontinued
//

Claims this corrects

Each of these is repeated widely, and each is wrong.
//

What we couldn't establish

A comparison that hides its gaps is less useful than one that names them.