3 July 2026
Getting text out of scans — in Bangla, English and 10 other languages
How OCR works, why Bangla OCR is harder than English, and how to prepare a scan so the text comes out clean.
A scanned page is a photograph. You can't search it, copy from it, or edit it until OCR (optical character recognition) turns the picture back into characters.
What to expect from OCR
Our OCR tool runs Tesseract in your browser — the document never leaves your device, which matters when the scan is a certificate, contract or ID. It reads 12 languages including বাংলা (Bangla), English, Arabic, Hindi, Chinese and Spanish.
Realistic accuracy expectations:
- Clean printed text, 300 DPI scan: near-perfect for English; very good for Bangla, with occasional trouble on conjunct characters (যুক্তাক্ষর) in low-quality scans.
- Photocopies of photocopies, faxes, phone photos at an angle: errors climb fast. The fix is a better input, not a different engine.
- Handwriting: classical OCR barely works at all. Use the separate Handwriting OCR tool, which uses an AI vision model instead — much better on handwriting, but the page content is sent to the AI provider, so don't use it for documents that must stay private.
Preparing the scan (this is 80% of the result)
- Resolution: 300 DPI is the sweet spot. Below 200, characters lose the strokes OCR relies on.
- Straight and flat: skewed or curved lines wreck recognition. Photographing a document? Use Scan to PDF — its scanner mode auto-detects the page edges and deskews.
- Contrast: grey text on grey background is the enemy. The scanner filter boosts contrast automatically.
- One language at a time where possible: telling the OCR engine the right language dramatically improves results — pick Bangla for Bangla documents rather than hoping the English model copes.
After OCR
Expect to proofread. Even 99% accuracy means roughly one error every two lines, and numbers (account numbers, NID numbers) deserve special attention because a misread digit looks plausible.