ViTSTR-Transducer: Cross-Attention-Free Vision Transformer Transducer for Scene Text Recognition.

RNN-T autoregressive language model cross-attention scene text recognition (STR) vision transformer (ViT)

Journal

Journal of imaging
ISSN: 2313-433X
Titre abrégé: J Imaging
Pays: Switzerland
ID NLM: 101698819

Informations de publication

Date de publication:
13 Dec 2023
Historique:
received: 01 11 2023
revised: 02 12 2023
accepted: 07 12 2023
medline: 22 12 2023
pubmed: 22 12 2023
entrez: 22 12 2023
Statut: epublish

Résumé

Attention-based encoder-decoder scene text recognition (STR) architectures have been proven effective in recognizing text in the real world, thanks to their ability to learn an internal language model. Nevertheless, the cross-attention operation that is used to align visual and linguistic features during decoding is computationally expensive, especially in low-resource environments. To address this bottleneck, we propose a cross-attention-free STR framework that still learns a language model. The framework we propose is ViTSTR-Transducer, which draws inspiration from ViTSTR, a vision transformer (ViT)-based method designed for STR and the recurrent neural network transducer (RNN-T) initially introduced for speech recognition. The experimental results show that our ViTSTR-Transducer models outperform the baseline attention-based models in terms of the required decoding floating point operations (FLOPs) and latency while achieving a comparable level of recognition accuracy. Compared with the baseline context-free ViTSTR models, our proposed models achieve superior recognition accuracy. Furthermore, compared with the recent state-of-the-art (SOTA) methods, our proposed models deliver competitive results.

Identifiants

pubmed: 38132694
pii: jimaging9120276
doi: 10.3390/jimaging9120276
pii:
doi:

Types de publication

Journal Article

Langues

eng

Subventions

Organisme : JSPS Kakenhi
ID : 22H00540
Organisme : RUPP-OMU/HEIP
ID : N/A

Auteurs

Rina Buoy (R)

Department of Core Informatics, Graduate School of Informatics, Osaka Metropolitan University, Osaka 599-8531, Japan.

Masakazu Iwamura (M)

Department of Core Informatics, Graduate School of Informatics, Osaka Metropolitan University, Osaka 599-8531, Japan.

Sovila Srun (S)

Department of Information Technology Engineering, Faculty of Engineering, Royal University of Phnom Penh, Phnom Penh 12156, Cambodia.

Koichi Kise (K)

Department of Core Informatics, Graduate School of Informatics, Osaka Metropolitan University, Osaka 599-8531, Japan.

Classifications MeSH