deftools.io Media

🎤 Audio Transcriber

Transcribe audio to text with SRT/VTT export. Speech recognition runs locally with Whisper.

Loading Whisper engine…

About this tool

Transcribe audio files or live microphone recordings into text. The transcription engine (Whisper) runs entirely in your browser — your audio never leaves your computer. Supports dozens of languages and can optionally translate speech to English.

How it works: select a model (tiny is fastest, small is most accurate), load an audio file or record from your mic, then click Transcribe. Results include timestamps that can be exported as SRT or VTT subtitle files for video editors and players.

The translate checkbox converts non-English speech directly into English text instead of transcribing in the source language.

FAQ

Why does the model need to download?

The Whisper speech recognition model is ~31 MB (tiny) to ~182 MB (small). It is downloaded once per session from HuggingFace and stored in browser memory. Page refreshes will re-download it.

Which model should I choose?

tiny (31 MB) is fastest — good for clear English speech with short audio. base (57 MB) balances speed and accuracy for most languages. small (182 MB) gives the best accuracy but takes longer to download and transcribe.

How long can my audio be?

Audio is limited to 120 seconds (2 minutes). Longer files are automatically truncated to the first 2 minutes.

What audio formats are supported?

Any format your browser can decode: WAV, MP3, MP4, OGG, WebM, FLAC, and more. Video files work too — only the audio track is processed.

Related media tools

Copied!