SVTR: Scene Text Recognition with a Single Visual Model-Reference-Cited by-同舟云学术

SVTR: Scene Text Recognition with a Single Visual Model

Published:2022-07 Issue: Volume: Page:
ISSN:
Container-title:Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence
language:
Short-container-title:

Author:

Du Yongkun¹,Chen Zhineng²,Jia Caiyan¹,Yin Xiaoting³,Zheng Tianlun²,Li Chenxia³,Du Yuning³,Jiang Yu-Gang²

Affiliation:

1. School of Computer and Information Technology and Beijing Key Lab of Traffic Data Analysis and Mining, Beijing Jiaotong University, China

2. Shanghai Collaborative Innovation Center of Intelligent Visual Computing, School of Computer Science, Fudan University, China

3. Baidu Inc., China

Abstract

Dominant scene text recognition models commonly contain two building blocks, a visual model for feature extraction and a sequence model for text transcription. This hybrid architecture, although accurate, is complex and less efficient. In this study, we propose a Single Visual model for Scene Text recognition within the patch-wise image tokenization framework, which dispenses with the sequential modeling entirely. The method, termed SVTR, firstly decomposes an image text into small patches named character components. Afterward, hierarchical stages are recurrently carried out by component-level mixing, merging and/or combining. Global and local mixing blocks are devised to perceive the inter-character and intra-character patterns, leading to a multi-grained character component perception. Thus, characters are recognized by a simple linear prediction. Experimental results on both English and Chinese scene text recognition tasks demonstrate the effectiveness of SVTR. SVTR-L (Large) achieves highly competitive accuracy in English and outperforms existing methods by a large margin in Chinese, while running faster. In addition, SVTR-T (Tiny) is an effective and much smaller model, which shows appealing speed at inference. The code is publicly available at https://github.com/PaddlePaddle/PaddleOCR.

Publisher

International Joint Conferences on Artificial Intelligence Organization

Cited by 75 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. SARN: Script-Aware Recognition Network for scene multilingual text recognition;Expert Systems with Applications;2024-09

2. Irregular text block recognition via decoupling visual, linguistic, and positional information;Pattern Recognition;2024-09

3. Scene text recognition: an Indic perspective;International Journal on Document Analysis and Recognition (IJDAR);2024-07-15

4. A Robust Pointer Meter Reading Recognition Method Based on TransUNet and Perspective Transformation Correction;Electronics;2024-06-21

5. Chitrantaran: Web-based Platform to Enhance the Document Digitization Process using OCR and Machine Translation;2024 4th Interdisciplinary Conference on Electrics and Computer (INTCEC);2024-06-11