Leveraging vision language models for digit recognition on digital measuring devices: a comparative analysis
Large language models (LLMs) with multimodal input capability have been demonstrated in various applications. Vision language models (VLMs), which were developed from LLMs, can process and learn both visual and textual data simultaneously, enabling them to generate descriptive text from images. This study explores the...