ViTSTR-Transducer: Cross-Attention-Free Vision Transformer Transducer for Scene Text Recognition.
RNN-T
autoregressive language model
cross-attention
scene text recognition (STR)
vision transformer (ViT)
Journal
Journal of imaging
ISSN: 2313-433X
Titre abrégé: J Imaging
Pays: Switzerland
ID NLM: 101698819
Informations de publication
Date de publication:
13 Dec 2023
13 Dec 2023
Historique:
received:
01
11
2023
revised:
02
12
2023
accepted:
07
12
2023
medline:
22
12
2023
pubmed:
22
12
2023
entrez:
22
12
2023
Statut:
epublish
Résumé
Attention-based encoder-decoder scene text recognition (STR) architectures have been proven effective in recognizing text in the real world, thanks to their ability to learn an internal language model. Nevertheless, the cross-attention operation that is used to align visual and linguistic features during decoding is computationally expensive, especially in low-resource environments. To address this bottleneck, we propose a cross-attention-free STR framework that still learns a language model. The framework we propose is ViTSTR-Transducer, which draws inspiration from ViTSTR, a vision transformer (ViT)-based method designed for STR and the recurrent neural network transducer (RNN-T) initially introduced for speech recognition. The experimental results show that our ViTSTR-Transducer models outperform the baseline attention-based models in terms of the required decoding floating point operations (FLOPs) and latency while achieving a comparable level of recognition accuracy. Compared with the baseline context-free ViTSTR models, our proposed models achieve superior recognition accuracy. Furthermore, compared with the recent state-of-the-art (SOTA) methods, our proposed models deliver competitive results.
Identifiants
pubmed: 38132694
pii: jimaging9120276
doi: 10.3390/jimaging9120276
pii:
doi:
Types de publication
Journal Article
Langues
eng
Subventions
Organisme : JSPS Kakenhi
ID : 22H00540
Organisme : RUPP-OMU/HEIP
ID : N/A