A BLSTM Network for Printed Bengali OCR System with High Accuracy

Debabrata Paul; Bidyut Baran Chaudhuri

arXiv:1908.08674·cs.CV·August 26, 2019·5 cites

A BLSTM Network for Printed Bengali OCR System with High Accuracy

Debabrata Paul, Bidyut Baran Chaudhuri

PDF

Open Access

TL;DR

This paper introduces a high-accuracy printed Bengali and English OCR system based on a simplified BLSTM-CTC architecture, achieving over 99% character accuracy across multiple fonts without using peephole connections or dropout.

Contribution

The paper presents a novel BLSTM-CTC OCR system for Bengali and English that omits peephole connections and dropout, resulting in improved accuracy and robustness.

Findings

01

Character accuracy of 99.32% on Bengali text

02

Word accuracy of 96.65% across 20 fonts

03

System is free and available online

Abstract

This paper presents a printed Bengali and English text OCR system developed by us using a single hidden BLSTM-CTC architecture having 128 units. Here, we did not use any peephole connection and dropout in the BLSTM, which helped us in getting better accuracy. This architecture was trained by 47,720 text lines that include English words also. When tested over 20 different Bengali fonts, it has produced character level accuracy of 99.32% and word level accuracy of 96.65%. A good Indic multi script OCR system is also developed by Google. It sometimes recognizes a character of Bengali into the same character of a non-Bengali script, especially Assamese, which has no distinction from Bengali, except for a few characters. For example, Bengali character for 'RA' is sometimes recognized as that of Assamese, mainly in conjunct consonant forms. Our OCR is free from such errors. This OCR system is…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsHandwritten Text Recognition Techniques · Natural Language Processing Techniques · Speech Recognition and Synthesis

MethodsDropout