Important
Superseded. This repository is the earlier (v2) MonOCR web app, kept for history. The web app now
lives in MonDevHub/monocr (apps/web), and the live site is
ocr.mondevhub.com. The model figures below describe the v2 model, not the
current one.
A privacy-first, in-browser OCR engine for the Mon language (mnw), powered by Rust, WebAssembly, and ONNX Runtime.
Note
The Mon language is classified as a "vulnerable" language in UNESCO's Atlas of the World’s Languages in Danger.
This project aims to digitize the Mon script, establishing a digital foundation suitable for future development, system integrations, and AI-driven preservation efforts.
MonOCR Web brings optical character recognition for the Mon script directly to the browser. By leveraging ONNX Runtime Web and a custom Wasm backend, all processing is performed locally on the user's device. Images are not sent to a server for recognition, and once the model is cached the app works offline. Files leave the browser only if the user opts in to Cloud Sync.
- On-Device Inference: Runs entirely in the browser via WebAssembly (Wasm).
- Local by Default: OCR runs on the device. Nothing is uploaded unless the user turns on Cloud Sync.
- Optional Cloud Sync: Secure, opt-in synchronization for contributing corrected scans to the open-source Mon language dataset.
- Model (v2, historical): MobileNetV3 + BiLSTM OCR engine (~6.6M parameters).
- Format Support: Handles PDFs and images up to 50MB.
- Script Specialized: Purpose-built for Mon script recognition, with supplementary support for Burmese and English.
Tip
File size is limited to 50MB for web and 20MB for mobile. For processing larger files or leveraging more powerful hardware, please use the CLI or package directly via uv add monocr or pip install monocr.
Image (Canvas/Blob)
LineSegmenter → horizontal projection profile → List<LineSegment>
ImagePreprocessor → grayscale + normalize [-1.0, 1.0]
MonOcrEngine → ONNX Runtime Web (monocr.onnx)
CtcDecoder → greedy CTC decode → String
This table describes the v2 model this repository ships, pinned to Hugging Face revision
a51be11 (src/lib/config.ts). For the current
model (v3.5, 11.55M parameters), see the MonOCR model card.
| Attribute | Specification |
|---|---|
| Architecture | MobileNetV3 + BiLSTM-384 + CTC |
| Precision | FP32 (ONNX) |
| Parameters | ~6.6M |
| Input | 128 × Variable (H × W) |
| Asset Size | ~25 MB |
monocr-web/
├── src/
│ ├── lib/
│ │ ├── monocr-onnx.ts # OCR Pipeline (ONNX/Wasm)
│ │ ├── components/ # Svelte UI Components
│ │ └── utils/ # Image & PDF Processing
│ └── routes/ # Application Pages
├── ocr-engine/ # Rust/Wasm engine (rten + ocrs)
├── functions/ # Cloudflare Pages model proxy
├── static/
│ ├── wasm/ # ONNX Runtime Wasm Binaries
│ └── fonts/ # Mon/Myanmar Unicode Fonts
└── scripts/ # Build & Asset Management
The web, Android and iOS apps now live together in MonDevHub/monocr:
- Web: live at ocr.mondevhub.com, built from
apps/web(this repository is its predecessor). - Android: native Jetpack Compose app in
apps/android. It is not in Google Play yet, so build it from source. - iOS: native SwiftUI app in
apps/ios. It is not in the App Store yet, so build it from source.
- Node.js 24+
- pnpm 11+
pnpm installCopy the pre-built ONNX Runtime WASM files to the static directory:
pnpm run copy-wasmpnpm devpnpm buildImportant
The build script automatically optimizes the monocr.onnx model deployment to comply with edge asset limits. In production, models are fetched from the HuggingFace CDN.
- HuggingFace Models (ONNX, Core ML)
- Unified SDKs
- NPM Package
- Help contribute to copy/translations here
MIT