Yangshu
AI-poweredFree

Scan in. Word, Excel or searchable PDF out.

OCRScribe — free AI OCR for documents

OCRScribe is a free Windows app that turns scanned contracts, invoices and photographed paperwork into editable Word files, searchable PDFs or Excel. Recognition runs on your own computer — nothing is uploaded.

OCRScribe, built by Yangshu, is a free AI OCR app for Windows that reads scanned documents offline in 109 languages and exports editable Word documents, searchable PDFs and Excel files.

Download for Windows

Windows 10/11 · 64-bit · Free · No sign-up · ≈ 137 MB installer (OCR models download during setup)

Use cases

Every document that only exists as a scan.

109 languages

It reads 109 languages, and works out which one on its own

Thai, English, Chinese, Japanese, Korean, Vietnamese, Indonesian, Malay, Arabic, Russian, German, French, Spanish, Portuguese and around ninety more — and mixing several on one page is the normal case, not an edge case. You never pick a language: OCRScribe detects it from the page itself. Drop in a PDF or photo, watch it read page by page, then export Word, searchable PDF or Excel.

Recognition and language detection are offline. The Excel and translation steps use AI — either the free daily allowance included with the app, or your own API key.

Among the 109

English中文ไทย日本語한국어Tiếng ViệtBahasa IndonesiaBahasa MelayuFilipinoភាសាខ្មែរລາວမြန်မာहिन्दीবাংলাதமிழ்اردوالعربيةفارسیעבריתРусскийУкраїнськаΕλληνικάTürkçeDeutschFrançaisEspañolPortuguêsItalianoNederlandsPolski+ 79 more

Contracts & agreements

Turn a signed contract from a flat scan into a Word file with its clauses, headings and tables intact — ready to amend, compare against a new draft, or hand to the AI for a translation.

Invoices & bills

Feed in a stack of paper invoices and get one Excel file back: OCR reads every page, the AI works out where one invoice ends and the next begins, and each invoice lands as its own row.

Cheques & drafts

Read a batch of cheques into an Excel sheet — number, amount, date, payee per row — in a shape you can upload straight into your accounting system instead of keying it in by hand.

Certificates & licences

Investment certificates, business licences and registration papers are dense, but you usually need only a few fields. Ask the AI for those and get a short table, not the whole page.

Foreign-language paperwork

Keep a Thai or Chinese contract word-for-word for legal validity, and have the AI produce a translated Word file alongside it for overseas counterparties to review.

Legacy scans & faxes

Faded print, stamps across the text, skewed phone photos and fax artefacts — switch to AI recognition for these, and anything still unclear is marked so you know where to look.

How it works

Four steps from a scan to a working file.

01

Drop in the file

Drag in a PDF or an image — JPG, PNG, TIFF, BMP, WEBP. Optionally tell the app what it is looking at: “invoice numbers are 13 digits”, “the first 12 pages are the contract”. That note measurably improves accuracy.

02

Watch it read

Recognized text appears on the left, the original page on the right, region by region, in the document's real reading order. Every page can be checked against its source as it goes.

03

Export what you need

Word for editing, searchable PDF for archiving, or Excel for the numbers. Word and PDF export are entirely local and need no API key.

04

Put the AI to work

Ask questions about the document, have it proofread or translated, or turn the whole batch into an Excel file. Every answer tells you which page it came from, so you can check it against the original.

Offline OCR that reads layout, not just characters

Two models run on your own machine: a layout model that finds every region on the page — titles, paragraphs, tables, figures, formulas, stamps — and returns their true reading order, and a vision model that reads each region. That is why the exported document keeps its structure instead of collapsing into one long column of text.

  • Nothing uploaded during local recognition
  • 25 region types with real reading order
  • Native-text PDFs skip OCR entirely — instant and exact

Three ways out: Word, searchable PDF, Excel

Word export rebuilds headings, tables and images as genuine Word objects, so the file is editable rather than a picture of a page. Searchable PDF keeps the scan looking exactly as it did, with every word selectable, searchable and copyable. And the AI builds the Excel or Word file for you, laid out around what your document actually contains.

  • Editable Word with real tables and images
  • Searchable PDF identical to the original scan
  • Excel and Word files produced by the AI

An AI assistant that has actually read the document

The assistant builds a memory of the whole file, then answers questions with page citations, tidies up the recognized text, translates it, or drafts a new document from it. A free daily allowance is included, so it works out of the box — or connect your own key for Gemini, OpenAI, Claude, DeepSeek or a compatible proxy.

  • Free daily AI allowance included — no key required
  • Bring your own key for unlimited use
  • Keys encrypted on your machine with Windows DPAPI

Uncertainty is flagged, never quietly guessed

Where the text can't be read clearly, OCRScribe marks the spot instead of guessing — so you can see exactly which parts need a second look. Flagged fields collect into a review list you can export to Excel, correct by hand and import back; the corrections then apply in bulk.

  • Low-confidence fields flagged with the reason
  • Export → correct in Excel → import back
  • Numbers, dates and reference codes are never rewritten

Built for real paperwork, and for real hardware

Phone photos are de-skewed and perspective-corrected, lighting is evened out, shadows removed, orientation fixed automatically. On the hardware side, OCRScribe benchmarks your machine once at first launch and picks the fastest device it finds — a discrete graphics card is roughly an order of magnitude faster than the CPU.

  • Auto perspective, orientation and lighting correction
  • Vulkan GPU acceleration — NVIDIA, AMD and Intel
  • 109 languages; Thai, Chinese and English mixed on one page

OCRScribe vs online OCR services

Why local recognition wins for business documents.

OCRScribeOnline OCR services
Documents uploaded to a serverNever during local recognitionRequired
Priced per page or per creditNo — completely freeUsually
Layout and reading order preservedBuilt in — 25 region typesOften flattened to plain text
Editable Word + searchable PDF + ExcelAll three, from one passUsually one format, often paid
Uncertain fields flagged for reviewYes, with a correction round-tripRare
Thai, Chinese and English on one pageHandled as the normal caseOften weak

FAQ

Common questions about OCRScribe.

Is it really free?

Yes — OCRScribe is completely free, with no account and no trial period. It is a tool Yangshu builds to demonstrate our AI engineering. A free daily AI allowance is included as well; beyond that you can connect your own API key, which is billed by that provider rather than by us.

Are my documents uploaded anywhere?

Local OCR runs entirely on your computer and uploads nothing at all. Text is sent to the AI provider you configured only when you choose to use an AI feature — questions, translation, Excel generation. If you switch the recognition method to AI recognition, page images are sent as well; that choice is always explicit.

Do I need an internet connection?

Only during installation. The installer downloads the OCR models — roughly 1.9 GB, once — after which recognition, Word export and searchable-PDF export all work with no connection at all. The AI features need internet because they call a cloud model.

What files can it read?

PDF and images — JPG, PNG, TIFF, BMP and WEBP. PDFs that already contain real text skip OCR for those pages, which is both instant and exact. A single PDF can be up to 200 pages; split anything longer and import the parts separately.

Which languages are supported?

109 in total. The most commonly used are English, Chinese (Simplified and Traditional), Japanese, Korean, Thai, Vietnamese, Indonesian, Malay, Filipino, Khmer, Lao, Burmese, Hindi, Bengali, Tamil, Urdu, Arabic, Persian, Hebrew, Russian, Ukrainian, Greek, Turkish, German, French, Spanish, Portuguese, Italian, Dutch and Polish. Several languages on one page are handled together, and the language is detected automatically — you never have to pick one. For the complete list, see the PaddleOCR-VL 1.6 open-source model that powers recognition.

What are the PC requirements?

Any 64-bit Windows 10 or 11 PC. It runs on the CPU without a graphics card, but a discrete GPU is roughly an order of magnitude faster — about 2.5 seconds per page against about 43 seconds on a CPU. OCRScribe benchmarks your machine at first launch and selects the fastest device automatically. The installer is about 137 MB and the models about 1.9 GB.

How accurate is it on Thai documents?

Thai is a first-class language here, not an afterthought — Thai, Chinese and English mixed on a single page is the normal case. Accuracy still depends on the source image: a flat, evenly lit, square-on scan reads far better than a shadowed photo. Anything the model is unsure about is flagged rather than guessed.

Can it handle a scan with stamps or handwriting?

Stamps, signatures and skewed scans are handled as first-class problems, and anything obscured is flagged for review. For heavy handwriting or badly damaged originals, switching the recognition method to AI recognition — where a cloud vision model reads each page image — usually gives a better result.

Is there a Mac version?

Not yet — OCRScribe is Windows 10/11 (64-bit) only for now.

Turn that pile of scans into files you can actually use.

Free, offline, no sign-up. Your documents never leave your computer.

Windows 10/11 · 64-bit · Free · No sign-up · ≈ 137 MB installer (OCR models download during setup)