A browser-based Optical Character Recognition (OCR) application for Kannada and other Indic languages. Built with Tesseract.js and Vue.js.
- Multi-language OCR — Supports 14 Indian languages, including Kannada + English mixed
- Default Kannada + English — Language dropdown defaults to
kan+engfor mixed-script documents - Styled Text Preservation — Bold, italic, and font sizes from the original document are preserved in the editor
- PDF Support — Upload and process multi-page PDF documents
- Image Support — Accepts JPG, PNG, GIF, BMP, TIFF formats
- Drag & Drop — Drag files directly or paste from clipboard
- Page Navigation — Navigate through PDF pages before processing
- Recognize All Pages — Batch process all pages in a PDF; editor shows page 1 immediately while remaining pages process in the background
- Per-Page Editing — OCR text is stored per page; navigate pages to view and edit each page's text individually
- Page View / Combined View — Toggle between per-page editing and a combined view of all pages
- Real-time Progress — See OCR progress for each page
- Rich Text Editor — Built-in TinyMCE editor with styled text (bold, italic, font sizes) preserved from OCR, proper paragraph and line break rendering for natural readability
- Word Diff — Track changes made to OCR'd text (added/removed words highlighted)
- Unique Words — Extract and copy unique words from OCR'd text
- Training Data Export — Export per-page PNG images,
.boxfiles (character-level bounding boxes), and.gt.txtground truth for Tesseract CLI training - Export Options — Export as TXT, DOCX, hOCR, HTML layout (with preserved styling), or TSV
- hOCR/HTML/TSV Viewer — View raw hOCR, HTML layout, and TSV data in a full-screen modal
- User Guide — Built-in help modal with usage instructions
- Server Storage — Optional: Store files and text on server for research
- Sanchaya Styling — Modern UI matching fonts.sanchaya.net
- 🆕 Crop & Copy Tool — Select specific regions of images and extract only the text from those areas
- Drag-to-select rectangular regions
- Undo/Redo with keyboard shortcuts (
Cmd+Z,Cmd+Shift+Z) - Copy extracted text to editor
- Multi-region support with smart state management
- Performance optimized for 60fps smooth drawing
| Language | Code | Language | Code |
|---|---|---|---|
| Assamese | asm | Bengali | ben |
| Gujarati | guj | Hindi | hin |
| Kannada | kan | Malayalam | mal |
| Marathi | mar | Odia | ori |
| Punjabi | pan | Sanskrit | san |
| Sinhala | sin | Tamil | tam |
| Telugu | tel | Urdu | urd |
| English | eng | Kannada+English | kan+eng |
- Website: https://ocr.sanchaya.net
- GitHub Pages: https://sanchaya.github.io/ocrsanchaya
The simplest setup - all processing happens in the browser.
# Clone the repository
git clone https://github.com/sanchaya/ocrsanchaya.git
cd ocrsanchaya
# Navigate to the built app
cd ocr-kannada/dist
# Start a local server
python3 -m http.server 8080Open http://localhost:8080 in your browser.
# Clone the repository
git clone https://github.com/sanchaya/ocrsanchaya.git
cd ocrsanchaya/ocr-kannada
# Install npm dependencies
npm install
# Start development server
npm run devOpen http://localhost:5173 in your browser.
cd ocrsanchaya/ocr-kannada
# Install dependencies and build
npm install
npm run build
# The built files will be in ocr-kannada/dist/
# Deploy the dist folder to any static hosting (Netlify, Vercel, etc.)This option stores uploaded files and OCR text on a server for research purposes.
- Python 3.11+
- Flask
# Clone the repository
git clone https://github.com/sanchaya/ocrsanchaya.git
cd ocrsanchaya
# Install Python dependencies
pip install -r requirements.txt
# Create required directories
mkdir -p uploads texts research
# Start the Flask server
python server.pyThe server will start on http://localhost:5001.
API Endpoints:
| Endpoint | Method | Description |
|---|---|---|
/api/upload |
POST | Upload image/PDF |
/api/save-text |
POST | Save OCR text |
/api/ocr-results |
GET | List OCR results |
/api/stats |
GET | Storage statistics |
/uploads/<filename> |
GET | Serve uploaded files |
/texts/<filename> |
GET | Serve text files |
To connect the frontend to the server:
Edit ocr-kannada/src/App.vue and set the server URL:
const SERVER_URL = 'https://ocr-server.sanchaya.net'; // or your server URLThen rebuild: npm run build and deploy.
# Build and run with Docker Compose
docker-compose up -d
# View logs
docker-compose logs -f
# Stop
docker-compose downServer runs on http://localhost:5001
Coolify is a self-hostable Heroku alternative.
- Log in to your Coolify dashboard
- Create a new project
-
Click Add New Resource → Docker
-
Configure:
- Repository:
https://github.com/sanchaya/ocrsanchaya - Branch:
master - Port:
5001
- Repository:
-
Save the resource
Add these persistent volumes to preserve data:
| Host Path | Container Path |
|---|---|
/app/uploads |
/app/uploads |
/app/texts |
/app/texts |
/app/research |
/app/research |
Or use the included coolify.json configuration:
- The
coolify.jsonfile in this repo will auto-configure these settings when you import it.
- Click Deploy
- Wait for the build to complete
- Your OCR server is now running!
Simply push new code to GitHub and redeploy in Coolify:
git add -A
git commit -m "Update"
git push origin masterThen click Redeploy in Coolify.
- Upload Image/PDF — Drag & drop, paste from clipboard, or click to select file
- Select Language — Choose from the dropdown (default: Kannada + English)
- Recognize — Click "Recognize" for single page, or "Recognize All Pages" for PDFs
- Edit Text — Use the built-in editor to correct OCR errors; original styling (bold, italic, font size) is preserved
- Navigate Pages (PDFs) — Use Prev/Next to switch pages; editor updates to show each page's OCR text
- View Modes — Toggle between Page View (per-page editing) and Combined View (all pages concatenated)
- Track Changes — See word diff when editing text
- Extract Words — Click "Unique Words" to get all unique words
- Training Data — Click "Train Data" to export
<page>.png+<page>.box+<page>.gt.txtfor Tesseract training - Export — Save as TXT, DOCX, hOCR, HTML (with styling), or TSV
- Help — Click ಸಹಾಯ | Help in the header for the built-in usage guide
- Server Storage — If configured, files and text are saved to the server
Extract text from specific regions of images using the new Crop & Copy tool.
- Upload image and run OCR recognition
- Click "Crop Mode" button to enable region selection
- Drag on image to select a rectangular region
- Click "Copy Region X" button to extract text from that area
- Text is automatically appended to the editor
| Shortcut | Action |
|---|---|
Cmd+Z (Mac) / Ctrl+Z |
Undo last selection |
Cmd+Shift+Z / Ctrl+Shift+Z |
Redo selection |
Esc |
Exit crop mode |
- Undo/Redo — Full history with 20-level stack
- Multi-select — Create multiple regions and extract from each
- Keyboard Navigation — Fully accessible with keyboard shortcuts
- Performance — Optimized 60fps drawing with smart throttling
- Memory Efficient — Bounded history and smart cleanup
- Works in Both Versions — Static HTML and Vue versions supported
See CROP_FEATURE_GUIDE.md for:
- Detailed step-by-step guide
- Advanced features
- Troubleshooting
- Code examples
- Performance notes
- Browser compatibility
ocrsanchaya/
├── server.py # Flask server for file storage
├── requirements.txt # Python dependencies
├── Dockerfile # Docker image definition
├── docker-compose.yml # Docker Compose config
├── coolify.json # Coolify configuration
├── README.md # This file
├── DOCUMENTATION.md # Full codebase documentation
├── CROP_FEATURE_GUIDE.md # Crop & Copy feature guide
├── docs/
│ └── index.html # HTML documentation page
├── tests/ # Automated test suite
│ ├── run-tests.js # Standalone test runner (no dependencies)
│ ├── crop-tool.test.js # Full mocha test suite
│ └── package.json # Test configuration
├── ocr-kannada/
│ ├── src/
│ │ ├── App.vue # Main Vue component (OCR, editor, export)
│ │ └── components/
│ │ ├── CropTool.vue # Crop & Copy Vue component
│ │ └── ImageLoader.vue # Image/PDF loader
│ ├── public/
│ │ └── CNAME # Custom domain config
│ └── dist/ # Built production files
├── js/
│ ├── crop-tool.js # Crop & Copy JavaScript implementation
│ └── tesseract-ocr.js # OCR integration
├── style/
│ └── ocr.css # Styling including crop tool styles
└── research/ # OCR results storage (server)
└── ocr_results.json
# Install dependencies (if not already done)
cd tests
npm install
# Run tests with standalone runner (no external dependencies required)
node run-tests.js
# Or run with mocha (if npm dependencies installed)
npm test
# Run with watch mode
npm run test:watch
# Generate coverage report
npm run test:coverage-
Start development servers:
# Terminal 1: Vue version (port 3000) cd ocr-kannada npm run dev # Terminal 2: Static version (port 8001) python3 -m http.server 8001
-
Test static version:
- Visit http://localhost:8001/index.html
- Upload image and run OCR
- Use crop feature with keyboard shortcuts
-
Test Vue version:
- Visit http://localhost:3000
- Upload image and run OCR
- Use crop feature with undo/redo
When running the server, these directories are created automatically:
uploads/- Stores uploaded images/PDFstexts/- Stores generated OCR text filesresearch/- Stores metadata (JSON)
- Tesseract.js - OCR engine
- Vue.js 3 - Frontend framework
- TinyMCE - Rich text editor
- PDF.js - PDF rendering
- Flask - Python web server
- Docker - Containerization
- Swathanthra Malayalam Computing (SMC) - For the tesseract-ocr-web project
- Sanchaya - For the design system
- Tesseract OCR - OCR engine
- Sanchi Foundation - For supporting Indic language technology
MIT License