A voice product is three pieces of software, and only one of them is the model we measure. Recognition is the piece that fails first and fails quietly — it does not crash, it just returns the wrong words, and everything downstream reasons about them confidently. Here is every speech-to-text system that matters, on the axes that actually decide a project.
Surveyed, not benchmarked. Nothing on this page was measured by OpenGauntlet. What that means →
None of this is benchmarked here. OpenGauntlet measures Conversational Language Humanlikeness in text. This section is a sourced survey of the recognition market, not a trial. Where a claim could not be verified against a primary source, it says so rather than smoothing it over.
Word error rate is the right metric here — and that is the opposite of the rule on Voices, where WER correlates negatively with how good speech sounds. Recognition's job is getting the words right, so WER measures exactly the thing you care about. It is only ever quoted here with the test set that produced it: every vendor publishes a number, and they are measured on different audio.
Diarization is not included by default. Knowing who spoke is a separate capability from knowing what was said — frequently a paid add-on, often an entirely separate model, and sometimes simply absent. Every row states which, because assuming it comes free is the most common and most expensive mistake in this market.