OCR Studio

Turning Printed Eastern Nagari pages into Living Text

A great deal of a community’s written heritage survives only on paper: out-of-print books, old magazines, pamphlets, song collections, community newsletters and school primers. Some was typed years ago in fonts that today’s computers cannot read properly. These pages cannot be searched, copied or reused, and every year the paper grows more fragile. OCR Studio was built by the BhoomiTech Heritage and Development Foundation to turn such pages back into usable digital text.

What it is

OCR Studio is a free app that runs in any modern web browser and can be installed on a phone, tablet or laptop like a regular app. It reads a photograph or scanned PDF of a printed page and turns the words on it into editable text.

  • Opens photos and PDFs. Multi-page PDFs can be read page by page, or you can jump straight to a page.
  • Recognises Bengali, Assamese and English. These can be recognised separately or together on the same page.
  • Fixes difficult pages. Deskew straightens tilted scans. Perspective corrects pages photographed at an angle. Enhance and black-and-white help with faded print, yellowed paper and stains.
  • Lets you check and correct. The page and the text appear side by side, so the text can be proofread and corrected.
  • Records a reading. A reader can record themselves reading the page aloud, and the recording stays with the page.
Built for the field

After the first visit, the whole app works offline, including text recognition and its language models. A team can digitise a village library or a family’s collection of old books without any internet. Recognition is quick because the engine loads once, not for every page.

Private by design

Recognition happens entirely inside the browser on the user’s own device. Pages, text and recordings are never uploaded to a server. They leave the device only when the person holding them chooses to save or share them. A family’s letters, a village committee’s records or a community’s unpublished songs can be digitised without leaving the community’s hands.

A tool for languages written only in Eastern Nagari

Several languages of North East India, such as Bishnupriya Manipuri, are written only in the Eastern Nagari (Bengali–Assamese) script and have no text-recognition tool of their own. Because they share the script, OCR Studio’s Bengali and Assamese recognition can read their printed pages. Words specific to these languages may need correcting, and the side-by-side editor makes that quick. OCR Studio also helps these communities build a recognition model for their own language:

  • Training data from proofreading. Every proofread page can be exported as line images paired with the correct text. These pairs are exactly what is needed to train a recognition model.
  • Using the new model. Once a model has been trained for the language, it can be loaded into the app and used alongside the built-in ones.

Building a Unicode corpus for Indian languages

Everything OCR Studio produces is standard Unicode text, including pages printed from older typing fonts. Once proofread, this text adds to a text corpus: a large, clean collection of written language.

Natural language processing (NLP), the field behind machine translation, speech recognition, spell-checkers, keyboards and search, cannot support a language without such corpora. National programmes such as Bhashini depend on them, and smaller languages are usually left out because their written material has never been digitised.

OCR Studio is built for this work:
  • A collection that grows page by page. Proofread pages are gathered into a collection.
  • One export for the whole corpus. The collection exports as a single corpus with the combined text, each page’s image and source details, and the training lines.
  • Speech data from readings. A recorded reading paired with its text is the kind of speech-and-text data used to train speech technology.
Credit and permission, recorded

Each file can carry its source details: title, author, year, publisher, language, script, who digitised it, and its permission status. These details travel with every page in every export. Material a community wants to keep to itself can be marked “community use only”, and that mark is flagged clearly whenever the collection is exported.

What it can digitise
  • Books and periodicals: out-of-print books, magazines, souvenirs and literary journals.
  • Songs and poetry: printed collections of songs, poems, hymns and folk verse.
  • Community documents: council resolutions, society records, notices and newsletters.
  • Learning materials: textbooks, primers and word lists in the community’s language.
  • Older typed text: documents in pre-Unicode fonts, converted to modern Unicode.
  • Handwritten pages: automatic recognition does not reliably read handwriting yet. A built-in Eastern Nagari keyboard lets volunteers type what they read beside the zoomed-in page instead.
Ways to save the work
  • Text file: the corrected Unicode text, ready for any document, website or database.
  • Searchable PDF: the page image with a hidden text layer that can be searched and copied.
  • Page bundle: the page image, text, recording, source details and training lines in one file.
  • Collection export: a complete corpus of every page added, ready for archives and language programmes.
Room to grow
  • Whole-book processing. Pages are recognised one at a time. Processing a whole PDF in one go would save hours on long books.
  • Corrections in the searchable PDF. The PDF’s hidden text is the machine-recognised text, before corrections. Carrying the corrected text into it would make the PDF as accurate as the text file.
  • Easier model training. Training a model from the exported lines still needs technical help outside the app. A guided or shared training service would put this within reach of more communities.
  • A shared home for collections. Collections live on each device. A companion website plugin could gather contributions from many volunteers and publish the texts that are free to share.
  • Handwriting. Reading handwritten Eastern Nagari automatically would need new models, and the manual transcriptions made in the app are a first step toward them.
  • Keyboards for more languages. The typing keyboard currently follows Bishnupriya Manipuri spelling rules.
  • Translated menus. The app’s own menus are in English. Offering them in Assamese, Bengali and Bishnupriya Manipuri would make it easier for more community volunteers to use.

OCR Studio gives communities a simple, private way to move their printed words from fragile paper into the digital age, and to make sure their languages are part of India’s digital future.

More from the blog