Every row on the leaderboard already names a serving backend, because the same weights produce different results depending on what runs them — different quantizations supported, different hardware backends, different throughput under load. Here is every LLM inference engine we could verify against a primary source, on the axes that decide a project — including a real benchmark OpenGauntlet ran on its own hardware, not just vendor claims.
Surveyed, not benchmarked. Nothing on this page was measured by OpenGauntlet's judge pipeline.