VoiceScribe: Free AI Offline Speech-to-Text Software for Windows
One hour of audio takes 4–6 hours to write up by hand. VoiceScribe transcribes it offline, labels the speakers, and lets your own AI prompt finish the job.

VoiceScribe is a free Windows transcription application built by Yangshu (YANGSHU CO., LTD) that performs speech recognition and speaker diarization entirely offline on your own machine, turning video or audio into a transcript with speaker labels and timestamps, and exporting it as a formatted Word document or a subtitle file. The recognition and diarization models ship inside the installer and run locally — audio files are never uploaded, no account is required, no data is collected, and transcription works with the network disconnected.
Once the transcript exists, VoiceScribe hands it to an open AI stage: you supply your own API key (OpenAI, Gemini, Claude, and others), pick the model, and write your own prompt to decide what the transcript becomes — meeting minutes, a lecture summary, an interview write-up, multilingual subtitles, a content outline, or any format you want. You are not locked into a handful of preset templates.
It fits any situation where a long recording has to become usable text: internal meetings, courses and training, legal interviews, journalism, podcast and video production.
Yangshu is a custom enterprise software and IT systems developer based in Thailand with R&D centers in both Thailand and China. Our work spans industry-specific ERP, industrial optimization software, applied AI solutions, and RPA. VoiceScribe is our free flagship side project.
What problem does VoiceScribe solve?
It solves the problem of recordings taking forever to write up — a problem everyone has and almost nobody has fully solved — without asking you to hand over your audio.
The industry rule of thumb for manual transcription is a 4:1 to 6:1 ratio: one hour of audio takes an experienced transcriptionist four to six hours to turn into a usable transcript, and someone doing it themselves typically needs five to eight (see Rev and GMR Transcription). Multiple speakers, accents, and background noise stretch that ratio further. A two-hour meeting means an afternoon of manual work.
And meetings keep multiplying. Microsoft's 2025 Work Trend Index, based on aggregated Microsoft 365 signals, reports that 30% of meetings now span multiple time zones — up eight percentage points since 2021 — while meetings starting after 8 p.m. rose 16% year over year (Microsoft WorkLab). Earlier survey work found that 56% of respondents struggle to summarize what happened in a meeting, and 80% would like AI to handle meeting summaries and action items (Microsoft Work Trend Index). The demand is clear; the open question is how to meet it.
That shaped VoiceScribe's design goal: first turn the recording into a structured "script" — who spoke, at what minute, and what they said — then let AI reshape that script into whatever you actually need.
After transcription: what can a transcript become?
Whatever your prompt tells it to become. That is precisely why we left the AI stage open.
Most transcription tools bury a few fixed buttons in the interface — "Generate meeting minutes," "Summarize" — each backed by a hard-coded prompt. That works for a standard meeting, but real work is rarely standard. A lecture summary, the insight extraction from a user interview, an objection list from a sales call, follow-up points from a clinical conversation: each needs a completely different structure.
So VoiceScribe treats it this way: the transcript is the raw material, the prompt is the recipe, and you write the recipe.
| Your scenario | What the prompt asks for | Output |
|---|---|---|
| Team meeting | Group by agenda item; list decisions and action items with owners | Meeting minutes |
| Course / lecture | Extract a topic outline; flag key points and examples | Lecture summary, revision notes |
| User interview | Cluster pain points by theme, quoting the participant's own words | Research insights |
| Sales call | Pull out objections, budget signals, and committed next steps | Follow-up list |
| Podcast / video | Strip filler words, smooth spoken phrasing, translate | Multilingual subtitles, show notes |
| Content creation | Rewrite a spoken draft into an article | First draft |
The surrounding freedom matters too:
- Choose your model. OpenAI, Gemini, Claude, and others — matched to the difficulty and budget of the task. A cheap small model for a simple summary; a stronger one for a long, complex document.
- Choose your cost. Usage bills directly to your own AI provider. We are not a middle layer, we do not mark anything up, and we do not resell tokens.
- Reuse what works. Save a prompt once it's tuned and apply it to the next recording of the same kind.
And you can skip the key entirely. Speech recognition itself requires no internet connection and no API key.
Who uses it? Six typical users
| User | Typical recording | What they want out of it |
|---|---|---|
| Internal meetings / boards | Strategy discussions, budget reviews, project retrospectives | Decisions and action items grouped by topic |
| Education and training | Classes, lectures, internal training | Topic outlines, verbatim transcripts, revision material |
| Legal and compliance | Interviews, internal investigations | Timestamped transcripts that can be checked against the audio |
| User research / HR | In-depth interviews, hiring, performance conversations | Insights or assessment points grouped by theme |
| Journalists and researchers | Interviews, fieldwork, focus groups | Quotable transcripts that map back to the original audio |
| Content and training teams | Course videos, product demos, podcasts | Multilingual subtitles with accurate timing |
They have one thing in common: the recording is long, and what they need is not a wall of text but a finished, usable result.
How it works: four steps, plus one optional step
- Drop in a video or audio file. No format conversion needed.
- Pick a model. The "fast" model handles common 5 languages at speed; the "multilingual" model covers 31 languages.
- Wait while transcription and speaker diarization run locally. The progress bar gives a deliberately conservative estimate, and it calibrates to your machine the more you use it.
- Export a Word document or a subtitle file.
Up to this point: no account, no email, no internet connection.
Optional: add your own AI API key in settings, write a prompt, and have the model turn the transcript into minutes, a summary, or translated subtitles. The key is stored locally, and requests go from your computer straight to the provider you chose.
How good are the subtitles VoiceScribe produces?
Good, because we hold to one principle: AI handles language, and the program calculates the timing according to industry conventions.
Subtitle timing leaves very little room for error. The Netflix Timed Text Style Guide, for instance, requires each subtitle event to last at least 5/6 of a second (20 frames at 24fps) and no more than 7 seconds, with a minimum two-frame gap between consecutive events, a reading-speed ceiling of roughly 17 characters per second for adult content, and a limit of 42 characters per line across a maximum of two lines (Netflix Partner Help Center, subtitle timing guidelines). Precision at that level cannot be secured by telling a model "please don't change the timecodes."
What we found in testing: no matter how tightly the prompt is constrained, the moment a language model touches the numbers, digits get miscopied, timings get "helpfully" adjusted, and a dropped line shifts everything after it. That produced the design principle behind the feature:
Let AI do what it is good at, and take back what it is not.
- Given to AI: language. Removing filler words, smoothing fragmented speech, translating into another language.
- Withheld from AI: every piece of timing data. The model sees text and never sees a timecode. When each subtitle appears and how long it stays is computed by the program, following the conventions above.
That boundary has a practical upside: because the program owns the timeline, you can be aggressive with the text in your prompt — rewrite it, shorten it, change the register, switch languages — without ever breaking the subtitle structure. In the worst case you get subtitles whose timing is perfectly accurate and whose wording is merely unpolished.
The principle generalizes beyond subtitles. It matches what we see in enterprise projects: making AI genuinely useful is not about handing it everything, but about drawing the line — what goes to its language ability, and what stays with the determinism of code.
Why is VoiceScribe free?
Because it is a Yangshu showcase project, not a revenue project.
Yangshu's actual business is custom enterprise software: industry-specific ERP, industrial optimization software, applied AI, and RPA. Rather than buy advertising, we would rather build something people find genuinely useful — full functionality, no sign-up, no paywall. People who find it useful tend to remember who made it, and that works better for us than an ad.
The bring-your-own-key (BYOK) approach keeps this simple on your side too: we are not a token reseller, your usage bills directly to the AI provider at market rates, and you can switch models whenever you like. The software itself stays free.
Getting VoiceScribe
Free download: https://yangshu.io/en/tools/voicescribe Requirements: Windows 10 / 11, 64-bit. No registration. Transcription requires no internet connection.
FAQ
Q: Is VoiceScribe fully offline? A: Speech recognition and speaker diarization run entirely on your machine, audio files are never uploaded, and it works offline. The AI stage that follows — minutes, summaries, subtitle translation — needs an internet connection and your own API key.
Q: Why do I have to supply my own API key? A: So you can choose the model freely (OpenAI, Gemini, Claude, and others), match cost to the difficulty of the task, and switch whenever you want. We are not a middle layer, we do not mark up usage, and we do not restrict which models you can use.
Q: Can I use the software without an API key? A: Yes. Transcription, speaker diarization, Word export, and source-language subtitles with accurate timecodes all work without a key. Only the AI stage needs one.
Q: Can I customize the prompt? A: Yes — that is the point of the feature. The software ships with common templates (meeting minutes, lecture summary, subtitle translation), but you can rewrite them completely so the same transcript produces whatever format suits your work, and save the tuned prompt for reuse.
Q: How is this different from the transcription built into Word, Teams, or Zoom? A: Three things: transcription runs locally, the output is a deliverable document with speaker labels and timestamps, and the prompt and model in the AI stage are entirely yours to choose.
Q: Which languages does it support? A: Two recognition models are included: a "fast" model for high-speed transcription of common languages, and a "multilingual" model covering 31 languages. The full language list is on the download page.
Q: What is speaker diarization? A: Speaker diarization is the automatic determination of which speaker produced which segment of audio. The result reads like a script: Speaker 1 and Speaker 2, each with what they said and when. It does not identify people — it only separates distinct voices — and the labels can be renamed.
Q: What hardware do I need? Is a dedicated GPU required? A: No dedicated GPU is required; an ordinary office laptop is enough. The software deliberately yields some compute back to the operating system so the machine stays responsive, which means transcription takes longer on lower-spec hardware. Minimum requirements are listed on the download page.
Q: Why is it free? Will that change? A: VoiceScribe is a Yangshu showcase project; our revenue comes from custom enterprise software development. The current version is fully functional, free, and requires no registration. Costs for the AI stage are billed to your own API account.
Frequently Asked Questions
Is VoiceScribe fully offline?
Speech recognition and speaker diarization run entirely on your machine, audio files are never uploaded, and it works offline. The AI stage that follows — minutes, summaries, subtitle translation — needs an internet connection and your own API key.
Why do I have to use my own API key?
You can choose the model freely (OpenAI, Gemini, Claude, and others), match cost to the difficulty of the task, and switch whenever you want. We are not a middle layer, we do not mark up usage, and we do not restrict which models you can use.
Can I use the software without an API key?
Yes. Transcription, speaker diarization, Word export, and source-language subtitles with accurate timecodes all work without a key. Only the AI stage needs one.
Can I customize the prompt?
Yes — that is the point of the feature. The software ships with common templates (meeting minutes, lecture summary, subtitle translation), but you can rewrite them completely so the same transcript produces whatever format suits your work, and save the tuned prompt for reuse.
How is this different from the transcription built into Word, Teams, or Zoom?
Three things: transcription runs locally, the output is a deliverable document with speaker labels and timestamps, and the prompt and model in the AI stage are entirely yours to choose.
Which languages does it support?
A: Two recognition models are included: a "fast" model for high-speed transcription of common 5 languages, and a "multilingual" model covering 31 languages.
What is speaker diarization?
Speaker diarization is the automatic determination of which speaker produced which segment of audio. The result reads like a script: Speaker 1 and Speaker 2, each with what they said and when. It does not identify people — it only separates distinct voices — and the labels can be renamed.
What hardware do I need? Is a dedicated GPU required?
No dedicated GPU is required; an ordinary office laptop is enough. The software deliberately yields some compute back to the operating system so the machine stays responsive, which means transcription takes longer on lower-spec hardware. Minimum requirements are listed on the download page.
Why is it free? Will that change?
VoiceScribe is a Yangshu showcase project; our revenue comes from custom enterprise software development. The current version is fully functional, free, and requires no registration. Costs for the AI stage are billed to your own API account.