Speech to Text
Upload an audio file and get back the words as text. Runs on our own server — see the FAQ below for what that means for your audio.
About this tool
Upload an audio file, get back the transcribed text and the detected language — useful for notes, captions, or as the first step before translating what was said.
Questions about Speech to Text
Is my audio sent anywhere?+
Yes — like Text to Speech, this needs a brief trip to our own server, since that’s the only way to run the transcription. Nothing is stored — the audio and the result are both deleted immediately after you get your text back.
Does this cost anything to use?+
No — it runs on a free, open-source speech recognition engine we host ourselves, not a paid service.
Which languages does it understand?+
English, Spanish, French, German, Arabic, Chinese, Russian, Portuguese, and Japanese, with the language detected automatically by default. If you already know the language, selecting it directly is more reliable — this matters most for languages that sound similar to one another, where auto-detect can occasionally pick the wrong one.
How accurate is it?+
Good for most of the supported languages with clear audio, but it varies — languages with less available training data come out noticeably less accurate than well-resourced languages like English or Spanish.
What about audio from a video?+
Pull the audio out first with a video-to-audio tool, then upload that file here. This tool works on audio files directly, not video files.