Skip to content
Conference

VizAdapt: novel dataset development with a voice-interactive visual question answering system for blind and low vision individuals

Aug 2026 · International Conference on Computer Graphics and Virtuality · Vol 14315, pp. 143150G - 143150G-11 · 0 citations · 36 references
Engineering

Abstract

Blind and Low Vision Individuals (BLV) encounter significant difficulties in comprehending complex visual environments, while current assistive technologies typically lack speech-interactive reasoning capabilities. To address this gap, this study fine-tunes the Qwen2.5-Omni framework using Llamafactory based on our enhanced VizWiz dataset—VizAdapt (which includes speech-converted queries)—and integrates Text-to-Speech (TTS) technology to develop a voice-based environmental interaction system for BLV. This innovatively achieves a closed-loop "speech-to-speech" interaction mechanism. Furthermore, the system incorporates a voice prompt feature that actively provides shooting guidance when a user uploads a blurry photo that is difficult to recognize, assisting BLV in adjusting their shooting method to obtain a clearer image. Experimental results demonstrate that our model has achieved state-of-the-art performance across multiple metrics, including text generation (BLEU, METEOR, ROUGE-L), semantic understanding (BERTScore), and task accuracy. This system not only provides real-time, natural environmental perception for BLV but also achieves state-of-the-art performance in vision-speech interaction tasks through its innovative multimodal fusion architecture, significantly enhancing users' independent living and social participation capabilities.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.