Author

B. Kővári

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

Examining the Visual Capabilities of Multimodal Large Language Models for Automotive Applications

Automatic recognition and classification of vehicle damages is an important research direction in modern computer vision and artificial intelligence, playing an increasingly significant role in industrial and practical applications. Traditional computer vision-based approaches can recognize and classify objects with high accuracy; however, achieving task-specific performance typically requires large amounts of annotated data, time-consuming training or fine-tuning, and extensive parameter optimization. This process is not only resource- and cost-intensive but also limits the rapid adaptability of the technology. The aim of this research is to investigate how effectively the latest Multimodal Large Language Models (MLLMs) can recognize types of vehicle damage in a zero-shot setting, i.e., without fine-tuning, and to evaluate how their performance can be further improved through prompt engineering and fewshot prompting. MLLMs have the advantage of being able to provide multiple forms of information from a single query and supplement their outputs with natural-language explanations. In contrast, traditional models are generally designed to perform only one predefined task. Therefore, within the framework of this project, the performance of a fine-tuned YOLO-based computer vision model is compared with that of MLLMs in vehicle damage classification. This comparison highlights a modern, data-efficient approach that achieves competitive performance through prompt engineering and in-context learning, eliminating the need for additional model training and opening new directions for automotive applications.

Márk Mitrenga, B. Kővári, Péter Gáspár · 0 citations