General-purpose speech recognition for meetings, captions, and multilingual audio.
The speech directory
Find your next
speech tool.
One place to explore the voices,
ears, and interpreters for your agent.
A curated directory of independently available services. Humlet connections are being prepared.
About this directory10 tools to explore
Curated, not rankedTranscribe speech with timestamps, speaker labels, and audio event information.
Stream generated speech for conversational agents and spoken content.
Generate spoken audio from text, returned as a file or an audio stream.
Process recorded audio or stream live speech into your application.
Convert text into generated audio for spoken responses and narration.
Speech recognition with code-switching and speaker diarization capabilities.
Translate spoken audio into text or speech in supported languages.
Streaming speech recognition with turn detection for conversational agents.
Generate speech for voice agents, interactive experiences, and narration.
A starting point for your speech stack.
This is a curated directory, not an exhaustive catalog or performance ranking. Descriptions summarize public provider documentation reviewed on September 10, 2026. Model, language, price, and regional availability can change.
Provider names identify their services; they do not imply a partnership or an active Humlet integration. Follow the source links and evaluate with your own audio before choosing.