Deep neural network with attention model for scene text recognition

Shuohao Li; Min Tang; Qiang Guo; Jun Lei; Jun Zhang

doi:10.1049/iet-cvi.2016.0404

IET Computer Vision (Oct 2017)

Deep neural network with attention model for scene text recognition

Shuohao Li,
Min Tang,
Qiang Guo,
Jun Lei,
Jun Zhang

Affiliations

Shuohao Li: College of Information System and ManagementNational University of Defense TechnologyNo. 109, Deya RoadChangshaPeople's Republic of China
Min Tang: Department of Computing ScienceUniversity of Alberta116 Street & 85 AvenueEdmontonAlbertaCanada
Qiang Guo: College of Information System and ManagementNational University of Defense TechnologyNo. 109, Deya RoadChangshaPeople's Republic of China
Jun Lei: College of Information System and ManagementNational University of Defense TechnologyNo. 109, Deya RoadChangshaPeople's Republic of China
Jun Zhang: College of Information System and ManagementNational University of Defense TechnologyNo. 109, Deya RoadChangshaPeople's Republic of China

DOI: https://doi.org/10.1049/iet-cvi.2016.0404
Journal volume & issue: Vol. 11, no. 7
pp. 605 – 612

Abstract

Read online

The authors present a deep neural network (DNN) with attention model for scene text recognition. The proposed model does not require any segmentation of the input text image. The framework is inspired by the attention model presented recently for speech recognition and image captioning. In the proposed framework, feature extraction, feature attention and sequence recognition are integrated in a jointly trainable network. Compared with previous approaches, the following contributions are mainly made. (i) The attention model is applied into DNN to recognise scene text, and it can effectively solve the sequence recognition problem caused by variable length labels. (ii) Rigorous experiments are performed across a number of challenging benchmarks, including IIIT5K, SVT, ICDAR2003 and ICDAR2013 datasets. Results in experiments show that the proposed model is comparable or better than the state‐of‐the‐art methods. (iii) This model only contains 6.5 million parameters. Compared with other DNN models for scene text recognition, this model has the least number of parameters so far.

Published in IET Computer Vision

ISSN: 1751-9632 (Print); 1751-9640 (Online)
Publisher: Wiley
Country of publisher: United Kingdom
LCC subjects: Medicine: Medicine (General): Computer applications to medicine. Medical informatics; Science: Mathematics: Instruments and machines: Electronic computers. Computer science: Computer software
Website: https://ietresearch.onlinelibrary.wiley.com/journal/17519640

About the journal

Abstract

Keywords