Skip to content

About

Kannada OCR Tool for Web - Do a live OCR on your images and texts.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

68 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Kannada OCR | ಕನ್ನಡ ಓಸಿಆರ್

A browser-based Optical Character Recognition (OCR) application for Kannada and other Indic languages. Built with Tesseract.js and Vue.js.

Features

  • Multi-language OCR — Supports 14 Indian languages, including Kannada + English mixed
  • Default Kannada + English — Language dropdown defaults to kan+eng for mixed-script documents
  • Styled Text Preservation — Bold, italic, and font sizes from the original document are preserved in the editor
  • PDF Support — Upload and process multi-page PDF documents
  • Image Support — Accepts JPG, PNG, GIF, BMP, TIFF formats
  • Drag & Drop — Drag files directly or paste from clipboard
  • Page Navigation — Navigate through PDF pages before processing
  • Recognize All Pages — Batch process all pages in a PDF; editor shows page 1 immediately while remaining pages process in the background
  • Per-Page Editing — OCR text is stored per page; navigate pages to view and edit each page's text individually
  • Page View / Combined View — Toggle between per-page editing and a combined view of all pages
  • Real-time Progress — See OCR progress for each page
  • Rich Text Editor — Built-in TinyMCE editor with styled text (bold, italic, font sizes) preserved from OCR, proper paragraph and line break rendering for natural readability
  • Word Diff — Track changes made to OCR'd text (added/removed words highlighted)
  • Unique Words — Extract and copy unique words from OCR'd text
  • Training Data Export — Export per-page PNG images, .box files (character-level bounding boxes), and .gt.txt ground truth for Tesseract CLI training
  • Export Options — Export as TXT, DOCX, hOCR, HTML layout (with preserved styling), or TSV
  • hOCR/HTML/TSV Viewer — View raw hOCR, HTML layout, and TSV data in a full-screen modal
  • User Guide — Built-in help modal with usage instructions
  • Server Storage — Optional: Store files and text on server for research
  • Sanchaya Styling — Modern UI matching fonts.sanchaya.net
  • 🆕 Crop & Copy Tool — Select specific regions of images and extract only the text from those areas
    • Drag-to-select rectangular regions
    • Undo/Redo with keyboard shortcuts (Cmd+Z, Cmd+Shift+Z)
    • Copy extracted text to editor
    • Multi-region support with smart state management
    • Performance optimized for 60fps smooth drawing

Supported Languages

Language Code Language Code
Assamese asm Bengali ben
Gujarati guj Hindi hin
Kannada kan Malayalam mal
Marathi mar Odia ori
Punjabi pan Sanskrit san
Sinhala sin Tamil tam
Telugu tel Urdu urd
English eng Kannada+English kan+eng

Live Demo


Installation

Option 1: GitHub Pages (No Server Required)

The simplest setup - all processing happens in the browser.

# Clone the repository
git clone https://github.com/sanchaya/ocrsanchaya.git
cd ocrsanchaya

# Navigate to the built app
cd ocr-kannada/dist

# Start a local server
python3 -m http.server 8080

Open http://localhost:8080 in your browser.


Option 2: Local Development

# Clone the repository
git clone https://github.com/sanchaya/ocrsanchaya.git
cd ocrsanchaya/ocr-kannada

# Install npm dependencies
npm install

# Start development server
npm run dev

Open http://localhost:5173 in your browser.


Option 3: Production Build

cd ocrsanchaya/ocr-kannada

# Install dependencies and build
npm install
npm run build

# The built files will be in ocr-kannada/dist/
# Deploy the dist folder to any static hosting (Netlify, Vercel, etc.)

Option 4: Run with Server Storage

This option stores uploaded files and OCR text on a server for research purposes.

Prerequisites

  • Python 3.11+
  • Flask
# Clone the repository
git clone https://github.com/sanchaya/ocrsanchaya.git
cd ocrsanchaya

# Install Python dependencies
pip install -r requirements.txt

# Create required directories
mkdir -p uploads texts research

# Start the Flask server
python server.py

The server will start on http://localhost:5001.

API Endpoints:

Endpoint Method Description
/api/upload POST Upload image/PDF
/api/save-text POST Save OCR text
/api/ocr-results GET List OCR results
/api/stats GET Storage statistics
/uploads/<filename> GET Serve uploaded files
/texts/<filename> GET Serve text files

To connect the frontend to the server:

Edit ocr-kannada/src/App.vue and set the server URL:

const SERVER_URL = 'https://ocr-server.sanchaya.net'; // or your server URL

Then rebuild: npm run build and deploy.


Option 5: Docker (Local)

# Build and run with Docker Compose
docker-compose up -d

# View logs
docker-compose logs -f

# Stop
docker-compose down

Server runs on http://localhost:5001


Option 6: Coolify Deployment

Coolify is a self-hostable Heroku alternative.

Step 1: Create a New Project in Coolify

  1. Log in to your Coolify dashboard
  2. Create a new project

Step 2: Add a Docker Resource

  1. Click Add New Resource → Docker

  2. Configure:

    • Repository: https://github.com/sanchaya/ocrsanchaya
    • Branch: master
    • Port: 5001
  3. Save the resource

Step 3: Configure Volumes (Important!)

Add these persistent volumes to preserve data:

Host Path Container Path
/app/uploads /app/uploads
/app/texts /app/texts
/app/research /app/research

Or use the included coolify.json configuration:

  • The coolify.json file in this repo will auto-configure these settings when you import it.

Step 4: Deploy

  1. Click Deploy
  2. Wait for the build to complete
  3. Your OCR server is now running!

Updating the Server

Simply push new code to GitHub and redeploy in Coolify:

git add -A
git commit -m "Update"
git push origin master

Then click Redeploy in Coolify.


Usage

  1. Upload Image/PDF — Drag & drop, paste from clipboard, or click to select file
  2. Select Language — Choose from the dropdown (default: Kannada + English)
  3. Recognize — Click "Recognize" for single page, or "Recognize All Pages" for PDFs
  4. Edit Text — Use the built-in editor to correct OCR errors; original styling (bold, italic, font size) is preserved
  5. Navigate Pages (PDFs) — Use Prev/Next to switch pages; editor updates to show each page's OCR text
  6. View Modes — Toggle between Page View (per-page editing) and Combined View (all pages concatenated)
  7. Track Changes — See word diff when editing text
  8. Extract Words — Click "Unique Words" to get all unique words
  9. Training Data — Click "Train Data" to export <page>.png + <page>.box + <page>.gt.txt for Tesseract training
  10. Export — Save as TXT, DOCX, hOCR, HTML (with styling), or TSV
  11. Help — Click ಸಹಾಯ | Help in the header for the built-in usage guide
  12. Server Storage — If configured, files and text are saved to the server

🆕 Crop & Copy Feature

Extract text from specific regions of images using the new Crop & Copy tool.

Quick Start

  1. Upload image and run OCR recognition
  2. Click "Crop Mode" button to enable region selection
  3. Drag on image to select a rectangular region
  4. Click "Copy Region X" button to extract text from that area
  5. Text is automatically appended to the editor

Keyboard Shortcuts

Shortcut Action
Cmd+Z (Mac) / Ctrl+Z Undo last selection
Cmd+Shift+Z / Ctrl+Shift+Z Redo selection
Esc Exit crop mode

Features

  • Undo/Redo — Full history with 20-level stack
  • Multi-select — Create multiple regions and extract from each
  • Keyboard Navigation — Fully accessible with keyboard shortcuts
  • Performance — Optimized 60fps drawing with smart throttling
  • Memory Efficient — Bounded history and smart cleanup
  • Works in Both Versions — Static HTML and Vue versions supported

Full Documentation

See CROP_FEATURE_GUIDE.md for:

  • Detailed step-by-step guide
  • Advanced features
  • Troubleshooting
  • Code examples
  • Performance notes
  • Browser compatibility

Project Structure

ocrsanchaya/
├── server.py              # Flask server for file storage
├── requirements.txt       # Python dependencies
├── Dockerfile             # Docker image definition
├── docker-compose.yml     # Docker Compose config
├── coolify.json          # Coolify configuration
├── README.md              # This file
├── DOCUMENTATION.md       # Full codebase documentation
├── CROP_FEATURE_GUIDE.md  # Crop & Copy feature guide
├── docs/
│   └── index.html         # HTML documentation page
├── tests/                 # Automated test suite
│   ├── run-tests.js       # Standalone test runner (no dependencies)
│   ├── crop-tool.test.js  # Full mocha test suite
│   └── package.json       # Test configuration
├── ocr-kannada/
│   ├── src/
│   │   ├── App.vue        # Main Vue component (OCR, editor, export)
│   │   └── components/
│   │       ├── CropTool.vue    # Crop & Copy Vue component
│   │       └── ImageLoader.vue # Image/PDF loader
│   ├── public/
│   │   └── CNAME          # Custom domain config
│   └── dist/              # Built production files
├── js/
│   ├── crop-tool.js       # Crop & Copy JavaScript implementation
│   └── tesseract-ocr.js   # OCR integration
├── style/
│   └── ocr.css            # Styling including crop tool styles
└── research/              # OCR results storage (server)
    └── ocr_results.json

Testing

Run Automated Tests

# Install dependencies (if not already done)
cd tests
npm install

# Run tests with standalone runner (no external dependencies required)
node run-tests.js

# Or run with mocha (if npm dependencies installed)
npm test

# Run with watch mode
npm run test:watch

# Generate coverage report
npm run test:coverage

Manual Testing

  1. Start development servers:

    # Terminal 1: Vue version (port 3000)
    cd ocr-kannada
    npm run dev
    
    # Terminal 2: Static version (port 8001)
    python3 -m http.server 8001
  2. Test static version:

  3. Test Vue version:


Environment Variables

When running the server, these directories are created automatically:

  • uploads/ - Stores uploaded images/PDFs
  • texts/ - Stores generated OCR text files
  • research/ - Stores metadata (JSON)

Technologies


Credits


License

MIT License

About

Kannada OCR Tool for Web - Do a live OCR on your images and texts.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages