Title
SVTR: Scene Text Recognition with a Single Visual Model.
Abstract
Dominant scene text recognition models commonly contain two building blocks, a visual model for feature extraction and a sequence model for text transcription. This hybrid architecture, although accurate, is complex and less efficient. In this study, we propose a Single Visual model for Scene Text recognition within the patch-wise image tokenization framework, which dispenses with the sequential modeling entirely. The method, termed SVTR, firstly decomposes an image text into small patches named character components. Afterward, hierarchical stages are recurrently carried out by component-level mixing, merging and/or combining. Global and local mixing blocks are devised to perceive the inter-character and intra-character patterns, leading to a multi-grained character component perception. Thus, characters are recognized by a simple linear prediction. Experimental results on both English and Chinese scene text recognition tasks demonstrate the effectiveness of SVTR. SVTR-L (Large) achieves highly competitive accuracy in English and outperforms existing methods by a large margin in Chinese, while running faster. In addition, SVTR-T (Tiny) is an effective and much smaller model, which shows appealing speed at inference. The code is publicly available at https://github.com/PaddlePaddle/PaddleOCR.
Year
DOI
Venue
2022
10.24963/ijcai.2022/124
European Conference on Artificial Intelligence
Keywords
DocType
ISSN
Computer Vision: Recognition (object detection, categorization),Computer Vision: Scene analysis and understanding
Conference
IJCAI 2022
Citations 
PageRank 
References 
0
0.34
0
Authors
8
Name
Order
Citations
PageRank
Yongkun Du100.68
Zhineng Chen219225.29
Caiyan Jia38113.07
Xiaoting Yin400.68
Tianlun Zheng500.34
Chenxia Li600.68
Yuning Du700.34
Yu-Gang Jiang83071152.58