A voice product can hear perfectly and still feel broken, because it talks over you or waits too long after you stop. That is not the recognizer and it is not the model — it is a separate layer deciding when a speaker has actually finished. Here is every VAD model, semantic turn-detector, orchestrator, and commercial endpointing feature that matters, on the axes that decide whether a conversation feels natural.
Surveyed, not benchmarked. Nothing on this page was measured by OpenGauntlet.