Audio in. Word out. 100% offline.
VoiceScribe — AI-powered speech to text
VoiceScribe is a free Windows app that turns any recording or video into a formatted Word transcript — speakers separated automatically, and not a single file ever leaves your computer.
VoiceScribe, built by Yangshu, is a free, AI-powered Windows app that transcribes audio and video fully offline, separates speakers automatically, and exports a formatted Word document — supporting 31 languages, with files never uploaded.
Download for WindowsWindows 10/11 · 64-bit · Free · No sign-up · ≈ 445 MB (speech-to-text models download during install)
Use cases
From meeting rooms to movie nights.
Subtitles for film & TV
Watching a foreign show with no subtitles — a Japanese drama, say? Run the video through VoiceScribe to transcribe the speech with timestamps, then let AI translate it into your language — subtitles for something that never had them.
Transcription is offline; the AI translation step uses your own cloud API key (OpenAI, Claude, and others).
- 1Transcribe the speech
- 2AI translates it
- 3Subtitles in your language
Meeting minutes
Drop in a two-hour meeting recording and get a speaker-labeled Word document while you grab a coffee — instead of writing it up by hand.
Interviews & podcasts
Transcribe interviews and episodes with each speaker separated, ready to edit into an article or show notes.
Lectures & training
Turn recorded classes, courses, and training sessions into searchable text your students or staff can review.
Legal & consulting
Archive client interviews and consultations as time-stamped transcripts — fully offline, so confidential audio never leaves the room.
Cross-border teams
Transcribe multilingual meetings across Southeast Asia — Thai, Vietnamese, Indonesian — then summarize them with AI in the language your team reads.
Support & call QA
Transcribe customer-service or sales calls, separate agent and caller, then use AI to score quality or pull the key points — turning call archives into reviewable records.
How it works
Three steps from a recording to a document.
Drop in your file
Drag in any video or recording — MP4, MP3, WAV, M4A, and more. Nothing uploads; processing starts on your own machine.
Transcribe & separate speakers
VoiceScribe transcribes offline and automatically tells apart who spoke when. Audition each voice and name the roles.
Export Word or AI summary
One click exports a formatted Word document with timestamps and speaker names — or generate a clean meeting summary with AI.
Offline & private by design
Transcription is 100% local. Your meeting recordings, client interviews, and confidential audio never touch a server. If you add an API key for AI summaries, it's encrypted on your machine with Windows DPAPI.
- Zero-upload transcription
- Works with no internet
- API key encrypted locally (DPAPI)
Automatic speaker separation
VoiceScribe detects each speaker, lets you preview their voice, and name the roles — e.g. “Teacher / Client”. The exported document reads like a script: role, then what they said.
- Auto-detects each speaker
- Audition & name every voice
- Merge the same person across segments
31 languages — optimized for Southeast Asia
Two offline speech-to-text models: SenseVoice-Small (blazing fast — Chinese, Cantonese, English, Japanese, Korean) and Fun-ASR-MLT-Nano (31 languages, notably strong on Thai, Vietnamese, and Indonesian).
- Thai, Vietnamese, Indonesian optimized
- Auto language detection
- Mixed-language recordings handled
AI summaries with your own key
Bring your own API key — OpenAI, Claude, DeepSeek, Gemini, or any compatible proxy — to turn a raw transcript into polished meeting minutes or an interview summary. Streamed live, 40 output languages, Markdown export.
- Your key, your models
- 40 output languages
- Live streaming + Word/Markdown export
VoiceScribe vs online transcription
Why offline wins for real work.
| VoiceScribe | Online services | |
|---|---|---|
| Files uploaded to a server | Never — stays on your PC | Required |
| Priced per minute | No — completely free | Usually |
| Speaker separation | Built in, automatic | Often extra or missing |
| Works with no internet | Yes (transcription) | No |
| Southeast Asian languages | Optimized (Thai, Vietnamese, Indonesian…) | Often weak |
FAQ
Common questions about VoiceScribe.
Is it really free?
Yes — VoiceScribe is completely free. It's a tool Yangshu builds to showcase our AI work. The only optional paid part is the AI summary feature, which runs on your own API key.
Do I need an internet connection?
Only during installation — the installer downloads the two offline speech-to-text models, so setup requires an internet connection. Once installed, transcription, speaker separation, and Word export all run fully offline. The only other online feature is the optional AI summary, which uses your own API key.
What files can it read?
Video — MP4, MKV, MOV, AVI, and more — and audio — WAV, MP3, M4A, AAC, FLAC, and more.
What are the PC requirements?
Any 64-bit Windows 10 or 11 PC. It runs on the CPU, so no graphics card is needed. The installer is about 445 MB and downloads the two offline speech-to-text models during installation — down from 2.6 GB in earlier versions — so you'll need an internet connection for setup.
Does the AI summary cost money?
VoiceScribe doesn't charge for it, but it uses your own API key (OpenAI, Claude, DeepSeek, Gemini, or a compatible proxy), so any usage is billed by that provider. Every other feature works without a key.
Is there a Mac version?
Not yet — VoiceScribe is Windows 10/11 (64-bit) only for now.
Turn your next recording into a document.
Free, offline, no sign-up. Your files never leave your computer.
Windows 10/11 · 64-bit · Free · No sign-up · ≈ 445 MB (speech-to-text models download during install)