Repository logo
  • English
  • Español
  • Log In
    New user? Click here to register. Have you forgotten your password?
Repository logo
  • Communities & Collections
  • All
  • English
  • Español
  • Log In
    New user? Click here to register. Have you forgotten your password?
  1. Home
  2. Browse by Author

Browsing by Author "Selvam, Prabu"

Now showing 1 - 1 of 1
Results Per Page
Sort Options
  • No Thumbnail Available
    Item
    A Transformer-Based Framework for Scene Text Recognition
    (Institute of Electrical and Electronics Engineers Inc., 2022) Selvam, Prabu; Sundar Koilraj, Joseph Abraham; Tavera Romero, Carlos Andres; Alharbi, Meshal; Mehbodniya, Abolfazl
    Scene Text Recognition (STR) has become a popular and long-standing research problem in computer vision communities. Almost all the existing approaches mainly adopt the connectionist temporal classification (CTC) technique. However, these existing approaches are not much effective for irregular STR. In this research article, we introduced a new encoder-decoder framework to identify both regular and irregular natural scene text, which is developed based on the transformer framework. The proposed framework is divided into four main modules: Image Transformation, Visual Feature Extraction (VFE), Encoder and Decoder. Firstly, we employ a Thin Plate Spline (TPS) transformation in the image transformation module to normalize the original input image to reduce the burden of subsequent feature extraction. Secondly, in the VFE module, we use ResNet as the Convolutional Neural Network (CNN) backbone to retrieve text image features maps from the rectified word image. However, the VFE module generates one-dimensional feature maps that are not suitable for locating a multi-oriented text on two-dimensional word images. We proposed 2D Positional Encoding (2DPE) to preserve the sequential information. Thirdly, the feature aggregation and feature transformation are carried out simultaneously in the encoder module. We replace the original scaled dot-product attention model as in the standard transformer framework with an Optimal Adaptive Threshold-based Self-Attention (OATSA) model to filter noisy information effectively and focus on the most contributive text regions. Finally, we introduce a new architectural level bi-directional decoding approach in the decoder module to generate a more accurate character sequence. Eventually, We evaluate the effectiveness and robustness of the proposed framework in both horizontal and arbitrary text recognition through extensive experiments on seven public benchmarks including IIIT5K-Words, SVT, ICDAR 2003, ICDAR 2013, ICDAR 2015, SVT-P and CUTE80 datasets. We also demonstrate that our proposed framework outperforms most of the existing approaches by a substantial margin.

Higher Education Institution subject to inspection and surveillance by the Ministry of National Education.
Legal status granted by the Ministry of Justice through Resolution No. 2,800 of September 2, 1959.
Recognized as a University by Decree No. 1297 of 1964 issued by the Ministry of National Education.

Institutionally Accredited in High Quality through Resolution No. 018144 of September 27, 2021, issued by the Ministry of National Education.

Ciudadela Pampalinda

Calle 5 # 62-00 Barrio Pampalinda
PBX: +57 (602) 518 3000
Santiago de Cali, Valle del Cauca
Colombia

Headquarters Centro

Carrera 8 # 8-17 Barrio Santa Rosa
PBX: +57 (602) 518 3000
Santiago de Cali, Valle del Cauca
Colombia

Palmira Section

Carrera 29 # 38-47 Barrio Alfonso López
PBX: +57 (602) 284 4006
Palmira, Valle del Cauca
Colombia

DSpace software copyright © 2002-2025 LYRASIS

  • Cookie settings
  • Privacy policy
  • End User Agreement
  • Send Feedback

Hosting & Support