A better way to compare multilingual speech.
Build an evaluation that tells you which system fits your users, not just which demo sounds impressive.
A short demonstration can make a capability easy to understand. Choosing infrastructure for a product requires another step: a repeatable evaluation using the conversations that product will encounter.
The useful question is specific. Which setup captures your customers’ speech, preserves the details, and responds within the time your application can tolerate?
Start with a representative set.
Collect consented examples of the languages, accents, vocabulary, and recording conditions you expect. Include spontaneous speech, changes of language, names, numbers, background noise, and corrections. Keep speakers used for tuning separate from evaluation speakers.
Write down what the set covers and what it leaves out. A small collection can reveal important problems, but it should not be presented as proof about every language or accent.
Give each system the same input.
Use identical recordings and document the model version, date, language settings, hints, and correction behavior. If one service requires a fixed language while another uses automatic detection, report the distinction rather than hiding it.
Compare tasks separately. Transcription quality, translated meaning, generated voice naturalness, and response delay are different measurements. A strong result in one does not establish a universal winner.
Measure meaning as well as text.
For transcription, native-speaker reference text supports word or character error measurement. Also inspect names, quantities, negation, and language-switch boundaries. A single aggregate score can hide a serious error in the detail your application uses.
For translation, use blind bilingual review with access to the original audio and context. For live systems, measure first-partial delay and the time from the end of speech to a completed output. Include long waits and failed requests in the report.
Test the automatic route itself.
If a system selects an engine automatically, replay the same recordings through that route and the relevant fixed alternatives. Compare quality, delay, and cost under the same constraints. Automatic selection should demonstrate a useful advantage for the intended application.
Humlet’s directory is a starting point for finding candidates. It is not a published performance ranking. Bring the evaluation set to the decision, show the cases that fail, and use the results to choose a setup you can explain.
A fair comparison ends with a decision for a defined use case, not a claim about every voice.
Building something that listens?
Explore the tools, or prepare a brief for your next speech integration.