A specialized Handwritten Digit Sequence Recognition (HDSR) framework for the Mexican Preliminary Election Results Program (PREP) based on a modified ResNet-18 architecture is proposed, introducing an asymmetric stride designed explicitly to preserve the 1:3 horizontal feature resolution of electoral tally sheets.
Abstract
The automated digitization of handwritten electoral results is critical for ensuring transparency and speed in democratic processes. While recurrent sequence-to-sequence models (e.g., CRNN+CTC) achieve high accuracy, they inherently violate the strict latency constraints of high-throughput administrative environments. Conversely, standard lightweight CNNs exhibit suboptimal performance on the long-tail distribution of high-cardinality scenarios. To bridge this gap, this study reformulates sequence recognition into a latency-bound classification task. We propose a specialized Handwritten Digit Sequence Recognition (HDSR) framework for the Mexican Preliminary Election Results Program (PREP) based on a modified ResNet-18 architecture. The methodology introduces an asymmetric stride designed explicitly to preserve the 1:3 horizontal feature resolution of electoral tally sheets, integrating a lightweight Convolutional Block Attention Module (CBAM) in deep stages to refine classification across 1001 possible sequences. Leveraging a megadiverse dataset of 3.77 million real-world images, the model was trained using AdamW and label smoothing to mitigate human-induced label noise. Results demonstrate a global accuracy of 97.82% and a significant improvement in Macro-Precision (0.8878) for rare sequences. With an inference latency of 9.1 ms on standard CPU hardware, the proposed solution offers a scalable, high-confidence alternative that prioritizes spatial preservation and fail-controlled deployment.
The proposed system outperforms several current CNN-, RNN-, and heuristic-based techniques, achieving character recognition accuracy of 97.89% and word recognition accuracy of 97.34%, and confirms the efficacy of integrating transformer-based learning with generative AI.
Padmavathi Pragada, D. Ch· Engineering Research Express· 0 citations
The authors introduce the Edge Suitability Score (ESS), a composite metric that combines normalized accuracy, model size, and inference speed into a single value, weighted at 0.40, 0.35, and 0.25 to reflect their relative importance for microcontroller deployment.
A compact convolutional network for 46-class DHCD Devanagari recognition and reached 99.73%, the highest reported at 15.6x smaller than prior state-of-the-art, effectively reaching the saturation point.
An edge-aware line-level HTR framework that extends a CNN-Transformer baseline with a learnable edge-extraction channel and Squeeze-and-Excitation channel attention and shows that combining learnable structural cues with channel-wise attention has improved robustness for degradation-prone historical manuscript collections.
Bilal Abdulrahman, Farhan Mohamed· Journal of Human Centered Te...· 0 citations
Handwritten text recognition (HTR) in examination scenarios has gained increasing attention for its role in intelligent grading systems. However, existing studies have not systematically modeled the complex handwriting phenomena inherent in exam settings, hindering a comprehensive understanding of the recognition challenges and limitations of current methods. Specifically, handwriting artifacts pose significant challenges to recognition models in two complementary aspects: sequentially, they disrupt the reading order and lead to non-monotonic sequences, while visually, they distort character structures and induce attention drift. To enable systematic benchmarking of exam handwriting, we first construct BNU-Exam-HTR, a large-scale dataset of handwritten exam text, and establish BNU-Exam-Benchmark, a fine-grained evaluation framework defining 12 representative challenges observed in real exam handwriting. To overcome these challenges, we further propose EduOCR, a recognition model with a collaborative dual-branch decoder. The Sequential Symbol Module (SSM) uses autoregressive decoding to handle non-monotonic sequences, while the Permutation-Aware Prediction Head (PPH) simulates artifact perturbations to guide the shared encoder in distinguishing characters from noise, thus stabilizing attention and mitigating alignment errors. Extensive experiments show that EduOCR consistently outperforms state-of-the-art HTR models, OCR tools, and multimodal large language models across all 12 challenges, demonstrating superior robustness and adaptability.
Runrui Li, Lin Zhu, Hua Huang· IEEE Transactions on Pattern...· 0 citations
Although exact word-level recognition remained difficult because of the extremely limited dataset size, the proposed framework successfully learned meaningful sequential patterns and produced increasingly structured Kirundi-like predictions.
Niyifasha Patrick· International journal of re...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.