The proposed system outperforms several current CNN-, RNN-, and heuristic-based techniques, achieving character recognition accuracy of 97.89% and word recognition accuracy of 97.34%, and confirms the efficacy of integrating transformer-based learning with generative AI.
Abstract
This study presents an advanced framework for Telugu handwritten character recognition by integrating Conditional Generative Adversarial Networks (cGANs) with Vision Transformer (ViT) architectures. Critical issues with Telugu scripts, such as intricate character structures, significant inter-writer variability, and a lack of annotated handwritten data, are addressed by the suggested method. While the Vision Transformer utilizes self-attention mechanisms to capture long-range spatial dependencies and global contextual features necessary for accurate recognition, cGAN-based synthetic data augmentation is employed to enhance dataset diversity and mitigate class imbalance. The proposed system outperforms several current CNN-, RNN-, and heuristic-based techniques, achieving character recognition accuracy of 97.89% and word recognition accuracy of 97.34%, as determined through extensive experiments conducted on real and synthetic handwritten datasets. Stable performance under noisy and real-world conditions is further confirmed by robustness analysis. The outcomes confirm the efficacy of integrating transformer-based learning with generative AI, creating a dependable and scalable OCR solution for low-resource Indic scripts, such as Telugu.
Handwritten character recognition plays a crucial role in optical character recognition systems, particularly for low-resource and structurally complex scripts such as Tamil. Despite significant advances in deep learning, accurate recognition of handwritten Tamil characters remains challenging due to large variations in writing styles, complex character structures, and high computational requirements of existing models. Conventional approaches often suffer from limited detection accuracy, inefficient parameter tuning, and poor generalization in real-world scenarios. To address these challenges, this paper proposes a novel deep learning–based handwritten Tamil character recognition framework. The proposed model employs an Adaptive Single Shot Detector (ASSD) for efficient and accurate character localization, with its performance further enhanced through hyperparameter optimization using the Enhanced Running City Game Optimizer (ERCGO). For classification, a Multi-scale Dilated Graph Attention Network (MDGANet) is introduced to effectively capture spatial and structural relationships within handwritten Tamil characters. Extensive experimental evaluations conducted on benchmark datasets demonstrate that the proposed framework achieves superior recognition accuracy with reduced computational complexity compared to existing methods, highlighting its effectiveness and suitability for practical OCR applications.
K. Manoj, M. Iyapparaja· Intelligent Data Analysis· 0 citations
A unified multi-attention framework was developed that explicitly optimises structural handwritten character script recognition through an adaptive multi-phase learning-rate schedule, incorporating global and hierarchical local transformer operations (ViT and Swin Transformer), respectively.
P. M. Kakde, S. Gulhane· African Journal Of Applied R...· 0 citations
An edge-aware line-level HTR framework that extends a CNN-Transformer baseline with a learnable edge-extraction channel and Squeeze-and-Excitation channel attention and shows that combining learnable structural cues with channel-wise attention has improved robustness for degradation-prone historical manuscript collections.
Bilal Abdulrahman, Farhan Mohamed· Journal of Human Centered Te...· 0 citations
Recognition of Indic handwritten text is a difficult issue with the complex formations of characters, variability of graphemes, ambiguity of strokes, and significant differences between writers. This is particularly problematic in scripts (such as Tamil and Kannada) where the form of the handwritten words and the composition structure often are not regular. Although recent CNN-RNN and CNN-Transformer architecture have achieved encouraging results, they either pay much attention to local visual representation or global context representation and do not consider the structural relationship of handwritten patterns at an appropriate level. This work will offer a solution to this drawback by suggesting a Hybrid CNN 10 Capsule 10 Transformer network with Connectionist Temporal Classification (CTC) to perform end-to-end handwritten word recognition. The proposed framework uses CNN layers to extract local visual features, the capsule module to encode structural and compositional relationships, and the Transformer to learn long-range sequence dependencies to be correctly transcribed. The model is tested on Tamil and Kannada handwritten data to check the effectiveness of cross-scripts. The results of the experiment indicate that the given architecture has a test accuracy of 85.46 percent and a Character Error Rate (CER) of 0.0281 on Tamil and 88.70 percent and a Character Error Rate (CER) of 0.0175 on Kannada. These results indicate that the hybrid framework proposed enhances the performance of cross-script handwritten text recognition as compared to baseline architectures.
A S Manjunath, Umesh Dadadahalli Ramu, Madhusudan G et al.· International journal of com...· 0 citations
Offline handwritten Marathi character recognition is still kind of hard research problem because there is so much variability within the same class ,and between classes they can look a bit similar ,also the strokes are complex and different people write in their own style. A lot of CNN and Transformer like methods either do not really capture long range relationships well enough, or they end up being too heavy computationally, you know not so efficient. So in this paper we suggest a Hybrid Vision Mamba and Transformer (HVMT), framework for stronger offline handwritten Marathi character recognition. The HVMT idea combines the hierarchical feature extraction power of Vision Mamba, which uses selective state-space modeling, with the contextual representation learning of a smaller Transformer encoder, and inside that encoder we use Multi-Head Self-Attention. Experiments are done on the public MHCD_GIETV2 dataset, where handwritten Marathi characters are collected from writers in different age groups and with diverse writing styles. Before training the images are turned into grayscale, then normalized, resized, and also augmented, to help the model generalize better. The proposed HVMT is compared with CNN, ResNet-50, EfficientNet-B0, ConvNeXt-Tiny, Vision Transformer (ViT-B/16), Swin Transformer-Tiny, and Vision Mamba, all under the same experimental setup. Experimental results show that the proposed framework achieved accuracy 87.11% , precision 87.09% , recall 87.10% and F1-score 87.09% which is better than the compared architectures. At the same time it only uses 26.9 million parameters, 2.6 GFLOPs, and inference time 1.305 ms per image. In other words, the HVMT framework seems to strike a workable tradeoff between recognition precision and compute efficiency. Because of this it is a good fit for things like intelligent document analysis, handwritten document digitization , archival preservation , and several other Indic script recognition tasks and more.
S. Khandakhani, Sachikanta Dash, Sasmita Padhy et al.· Journal of Intelligent Decis...· 0 citations
This paper introduces NepScript Genesis, a Neural Architecture Search (NAS) framework for automated Generative Adversarial Network (GAN) discovery, applied to conditional Devanagari handwritten digit synthesis. We compare five NAS strategies against a carefully constructed Deep Convolutional GAN (DCGAN) baseline (FID=332.28). Architecture selection utilizes a two-stage pipeline guided by a novel domain-aware evaluation metric (Enhanced Score). Results demonstrate that Adaptive Exploration achieves the optimal quality-efficiency trade-off, attaining an FID of 79.12 -- a 76.19% improvement over the baseline -- and the highest mode coverage among the NAS strategies (Recall=0.531) in under one GPU-hour. Furthermore, we demonstrate that incorporating script-specific structural heuristics into the search phase prevents early-stage mode collapse. In a downstream low-resource evaluation, augmenting 250 real training samples per class with GAN-generated digits from the best NAS model improves CNN classification accuracy from 91.0% to 96.5% (+5.5 percentage points), demonstrating that NAS-optimized synthesis produces digits of sufficient quality to benefit practical recognition pipelines when real data is scarce.
Mausam Gurung, Prabin Neupane, S. Acharya· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.