Yangshu
AI-poweredFree

Audio in. Word out. 100% offline.

VoiceScribe — AI-powered speech to text

VoiceScribe is a free Windows app that turns any recording or video into a formatted Word transcript — speakers separated automatically, and not a single file ever leaves your computer.

VoiceScribe, built by Yangshu, is a free, AI-powered Windows app that transcribes audio and video fully offline, separates speakers automatically, and exports a formatted Word document — supporting 31 languages, with files never uploaded.

Download for Windows

Windows 10/11 · 64-bit · Free · No sign-up · ≈ 445 MB (speech-to-text models download during install)

Use cases

From meeting rooms to movie nights.

Popular

Subtitles for film & TV

Watching a foreign show with no subtitles — a Japanese drama, say? Run the video through VoiceScribe to transcribe the speech with timestamps, then let AI translate it into your language — subtitles for something that never had them.

Transcription is offline; the AI translation step uses your own cloud API key (OpenAI, Claude, and others).

  1. 1Transcribe the speech
  2. 2AI translates it
  3. 3Subtitles in your language

Meeting minutes

Drop in a two-hour meeting recording and get a speaker-labeled Word document while you grab a coffee — instead of writing it up by hand.

Interviews & podcasts

Transcribe interviews and episodes with each speaker separated, ready to edit into an article or show notes.

Lectures & training

Turn recorded classes, courses, and training sessions into searchable text your students or staff can review.

Legal & consulting

Archive client interviews and consultations as time-stamped transcripts — fully offline, so confidential audio never leaves the room.

Cross-border teams

Transcribe multilingual meetings across Southeast Asia — Thai, Vietnamese, Indonesian — then summarize them with AI in the language your team reads.

Support & call QA

Transcribe customer-service or sales calls, separate agent and caller, then use AI to score quality or pull the key points — turning call archives into reviewable records.

How it works

Three steps from a recording to a document.

01

Drop in your file

Drag in any video or recording — MP4, MP3, WAV, M4A, and more. Nothing uploads; processing starts on your own machine.

02

Transcribe & separate speakers

VoiceScribe transcribes offline and automatically tells apart who spoke when. Audition each voice and name the roles.

03

Export Word or AI summary

One click exports a formatted Word document with timestamps and speaker names — or generate a clean meeting summary with AI.

Offline & private by design

Transcription is 100% local. Your meeting recordings, client interviews, and confidential audio never touch a server. If you add an API key for AI summaries, it's encrypted on your machine with Windows DPAPI.

  • Zero-upload transcription
  • Works with no internet
  • API key encrypted locally (DPAPI)

Automatic speaker separation

VoiceScribe detects each speaker, lets you preview their voice, and name the roles — e.g. “Teacher / Client”. The exported document reads like a script: role, then what they said.

  • Auto-detects each speaker
  • Audition & name every voice
  • Merge the same person across segments

31 languages — optimized for Southeast Asia

Two offline speech-to-text models: SenseVoice-Small (blazing fast — Chinese, Cantonese, English, Japanese, Korean) and Fun-ASR-MLT-Nano (31 languages, notably strong on Thai, Vietnamese, and Indonesian).

  • Thai, Vietnamese, Indonesian optimized
  • Auto language detection
  • Mixed-language recordings handled

AI summaries with your own key

Bring your own API key — OpenAI, Claude, DeepSeek, Gemini, or any compatible proxy — to turn a raw transcript into polished meeting minutes or an interview summary. Streamed live, 40 output languages, Markdown export.

  • Your key, your models
  • 40 output languages
  • Live streaming + Word/Markdown export

VoiceScribe vs online transcription

Why offline wins for real work.

VoiceScribeOnline services
Files uploaded to a serverNever — stays on your PCRequired
Priced per minuteNo — completely freeUsually
Speaker separationBuilt in, automaticOften extra or missing
Works with no internetYes (transcription)No
Southeast Asian languagesOptimized (Thai, Vietnamese, Indonesian…)Often weak

FAQ

Common questions about VoiceScribe.

Is it really free?

Yes — VoiceScribe is completely free. It's a tool Yangshu builds to showcase our AI work. The only optional paid part is the AI summary feature, which runs on your own API key.

Do I need an internet connection?

Only during installation — the installer downloads the two offline speech-to-text models, so setup requires an internet connection. Once installed, transcription, speaker separation, and Word export all run fully offline. The only other online feature is the optional AI summary, which uses your own API key.

What files can it read?

Video — MP4, MKV, MOV, AVI, and more — and audio — WAV, MP3, M4A, AAC, FLAC, and more.

What are the PC requirements?

Any 64-bit Windows 10 or 11 PC. It runs on the CPU, so no graphics card is needed. The installer is about 445 MB and downloads the two offline speech-to-text models during installation — down from 2.6 GB in earlier versions — so you'll need an internet connection for setup.

Does the AI summary cost money?

VoiceScribe doesn't charge for it, but it uses your own API key (OpenAI, Claude, DeepSeek, Gemini, or a compatible proxy), so any usage is billed by that provider. Every other feature works without a key.

Is there a Mac version?

Not yet — VoiceScribe is Windows 10/11 (64-bit) only for now.

Turn your next recording into a document.

Free, offline, no sign-up. Your files never leave your computer.

Windows 10/11 · 64-bit · Free · No sign-up · ≈ 445 MB (speech-to-text models download during install)