en

Audio to text

Turn a spoken recording into written text - locally in the browser via AI speech recognition (Whisper), entirely without upload.

Running locally on your device ...

Recording language
  • Auto-detect
  • English
  • 中文
  • Deutsch
  • Español
  • Русский
  • 한국어
  • Français
  • 日本語
  • Português
  • Türkçe
  • Polski
  • Català
  • Nederlands
  • العربية
  • Svenska
  • Italiano
  • Bahasa Indonesia
  • हिन्दी
  • Suomi
  • Tiếng Việt
  • עברית
  • Українська
  • Ελληνικά
  • Bahasa Melayu
  • Čeština
  • Română
  • Dansk
  • Magyar
  • தமிழ்
  • Norsk
  • ไทย
  • اردو
  • Hrvatski
  • Български
  • Lietuvių
  • Latina
  • Māori
  • മലയാളം
  • Cymraeg
  • Slovenčina
  • తెలుగు
  • فارسی
  • Latviešu
  • বাংলা
  • Српски
  • Azərbaycanca
  • Slovenščina
  • ಕನ್ನಡ
  • Eesti
  • Македонски
  • Brezhoneg
  • Euskara
  • Íslenska
  • Հայերեն
  • नेपाली
  • Монгол
  • Bosanski
  • Қазақша
  • Shqip
  • Kiswahili
  • Galego
  • मराठी
  • ਪੰਜਾਬੀ
  • සිංහල
  • ខ្មែរ
  • ChiShona
  • Yorùbá
  • Soomaali
  • Afrikaans
  • Occitan
  • ქართული
  • Беларуская
  • Тоҷикӣ
  • سنڌي
  • ગુજરાતી
  • አማርኛ
  • ייִדיש
  • ລາວ
  • Oʻzbekcha
  • Føroyskt
  • Kreyòl ayisyen
  • پښتو
  • Türkmençe
  • Nynorsk
  • Malti
  • संस्कृतम्
  • Lëtzebuergesch
  • မြန်မာ
  • བོད་སྐད་
  • Tagalog
  • Malagasy
  • অসমীয়া
  • Татарча
  • ʻŌlelo Hawaiʻi
  • Lingála
  • Hausa
  • Башҡортса
  • Basa Jawa
  • Basa Sunda
  • 粵語

Running locally on your device ...

0%

Is my file uploaded?

No. Everything runs in your browser - your file never leaves your device. How this is verifiable

No upload100% local
Your content stays with youno third-party access
Hosted in GermanyGlobal Content Delivery
Independently auditedTLS A+ · HTTP headers A+

An interview, a lecture recording, a voice memo capturing an idea on the go: this tool listens to the recording and writes it down as searchable, copyable text. The recognition model is OpenAI's Whisper, here as a purpose-compiled, single-thread WebAssembly build - the same underlying technology well-known transcription services use, except it runs entirely on your device instead of a foreign server.

By default, Whisper detects the spoken language itself ("Auto-detect") - for short or multilingual recordings, the language can also be fixed explicitly, which can slightly improve accuracy and speed. The result appears as running text with a paragraph per speech segment; for videos that need embedded timestamps, the sibling tool "Video to subtitles" uses the same recognition but outputs an SRT/VTT file with exact time marks instead.

Because there is no shared memory (no cross-origin isolation), recognition runs single-threaded rather than multi-threaded - correct, but noticeably slower than a server-side service. That is why a much stricter length cap applies here than for the site's other audio tools; honestly slow beats promising a false speed.

Specifications

Specifications
Input formatsMP3, WAV, M4A, AAC, OGG, OPUS, FLAC
Auto-conversionMP4
Output formatTXT
Batch processingNo
ProcessingLocally in your browser (WebAssembly)
File uploadNone

In 3 steps

  1. Drop your audio file.
  2. Pick a language or leave it on "Auto-detect".
  3. Copy the recognised text or download it as TXT.

Limitations: Runs single-threaded (no cross-origin-isolation mode on this site) and is therefore noticeably slower than server-side services - a much stricter per-device length cap applies here (shorter than the other audio tools). Uses the compact "tiny" model tier: good for clear speech, weaker with heavy background noise, dialects or jargon than a larger model would be. Recognition is not error-free - check the result, especially proper names and numbers.

FAQ

Is my recording uploaded?

No, recognition runs entirely locally in the browser via WebAssembly.

Which languages are recognised?

Nearly 100 languages; either auto-detected or explicitly chosen.

Why does this take longer than other audio tools?

Speech recognition is compute-heavy and runs here without multi-threading - correct, but deliberately slower rather than falsely promised as fast.

Is there a version with timestamps for subtitles?

Yes, the sibling tool "Video to subtitles" outputs SRT/VTT with exact time marks.

Related tools

Video to subtitles · Remove noise · Image to text (OCR) · Create a waveform image · Record audio · Swap left/right