Oct 2026· Bulletin of Electrical Engineering and Informatics· 0 citations· 29 references
Handwritten Text Recognition Techniques
Abstract
Large language models (LLMs) with multimodal input capability have been demonstrated in various applications. Vision language models (VLMs), which were developed from LLMs, can process and learn both visual and textual data simultaneously, enabling them to generate descriptive text from images. This study explores the use of several VLM models for the seven-segment digit recognition of a digital measuring device compared with other baseline optical character recognition (OCR) methods, such as Tesseract OCR, EasyOCR, PaddleOCR, KerasOCR, and YOLO11, with or without the integration of you only look once 11 (YOLO11) for LCD screen cropping. The results showed that the Gemini model outperformed conventional OCR methods and other VLM models with an average per-digit-recognition (PDR) accuracy of 98.10% (?=0.36%) and full sequence reading (FSR) accuracy of 93.85% (?=1.29%). The performance evaluation across the distance variation also demonstrates high accuracy for the VLM result with stable results. Simultaneously, YOLO11 exhibits degradation at longer distances due to difficulty in detecting small decimal digits, whereas the traditional OCR method yields inconsistent results. This study showed promising results for the further development, improvement, and deployment of VLM models in the automatic seven-segment digit recognition task on the edge device.
Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...
O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al.· Zenodo (CERN European Organi...· 465 citations
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.
P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al.· IEEE International Conferenc...· 110 citations· ⚡7