//opengauntlet
// surveyed, not benchmarked — 35 systems

Voice agent turn detection and VAD compared

A voice product can hear perfectly and still feel broken, because it talks over you or waits too long after you stop. That is not the recognizer and it is not the model — it is a separate layer deciding when a speaker has actually finished. Here is every VAD model, semantic turn-detector, orchestrator, and commercial endpointing feature that matters, on the axes that decide whether a conversation feels natural.

Surveyed, not benchmarked. Nothing on this page was measured by OpenGauntlet.

How to read this section

None of this is benchmarked here. OpenGauntlet measures Conversational Language Humanlikeness in text. This section is a sourced survey of the turn-taking market, not a trial. Where a claim could not be verified against a primary source, it says so rather than smoothing it over.

Four kinds of thing share this page because a builder chooses between them as alternatives even though they are architecturally different: a raw VAD model that decides speech-vs-silence, a semantic turn-detector that decides whose turn it is, an orchestrator framework that bundles one of the above, or a commercial API offering endpointing as a feature. Orchestrator rows describe ONLY their turn-taking behaviour — not their transport, tool-calling, or general architecture.

//

Every system, compared

Filter by kind or by what a system can actually do, pick the columns you care about, search any field, click a heading to sort. Every row carries the one thing a buyer would otherwise find out too late.
Loading…
commercial-safe weights permissive, with a catch not usable commercially proprietary service discontinued
//

Claims this corrects

Each of these is repeated widely, and each is wrong.
//

What we couldn't establish

A comparison that hides its gaps is less useful than one that names them.