MediaMintMediaMint

Speech to Text

Upload an audio file and get back the words as text. Runs on our own server — see the FAQ below for what that means for your audio.

About this tool

Upload an audio file, get back the transcribed text and the detected language — useful for notes, captions, or as the first step before translating what was said.

Questions about Speech to Text

Is my audio sent anywhere?+

Yes — like Text to Speech, this needs a brief trip to our own server, since that’s the only way to run the transcription. Nothing is stored — the audio and the result are both deleted immediately after you get your text back.

Does this cost anything to use?+

No — it runs on a free, open-source speech recognition engine we host ourselves, not a paid service.

Which languages does it understand?+

English, Spanish, French, German, Arabic, Chinese, Russian, Portuguese, and Japanese, with the language detected automatically by default. If you already know the language, selecting it directly is more reliable — this matters most for languages that sound similar to one another, where auto-detect can occasionally pick the wrong one.

How accurate is it?+

Good for most of the supported languages with clear audio, but it varies — languages with less available training data come out noticeably less accurate than well-resourced languages like English or Spanish.

What about audio from a video?+

Pull the audio out first with a video-to-audio tool, then upload that file here. This tool works on audio files directly, not video files.

More tools