Skip to content

A Multi-Modal Deep Learning Framework for Assistive Vision: Software-Based Design and Real-Time Validation

· 0 citations · 14 references

TL;DR

The development of an intelligent visual aid incorporating object detection, face recognition, distance estimation, distance estimation, and text-to-speech audio feedback is outlined.

View source

Similar papers

Conference Aug 2026

Smart Sight: A Comprehensive Deep Learning Framework for Real-Time Assistive Navigation and Object Recognition for Visually Impaired Individuals

Visual impairment is a universal health burden, and one of the most pressing concerns in the world, as it is estimated that there are 2.2 billion afflicted persons in the world, which is highly constraining in their free movements, their interactions with the surrounding, as well as on their social interrelation. White...

S. M. Raj, V. Hemanth, N. Nikhil et al. · 0 citations

AI-Based Object Detection and Smart Navigation System for the Visually Impaired

An AI-powered assistive system designed to enhance the mobility, safety, and independence of visually impaired individuals, creating a smart, voice-guided companion that empowers visually impaired users to navigate their surroundings with confidence and independence.

G. Sireesha, K. Prasanthi, Kanchumarthi Nirmala et al. · 0 citations
Open access Sep 2026

Integrating Computer Vision and Large Language Models for Real-Time Blind Navigation

The emergence of artificial intelligence (AI) in assistive technology is a rapidly evolving field that offers a great opportunity for visually impaired people to gain greater freedom. Despite the high pace of development in AI, current solutions are still rather incoherent and expensive and lack contextual reasoning an...

Ashish Sharma, Dhiraj Rajput, Sushant Kumar et al. · 0 citations
Open access 2026

Vision-Enabled Virtual Robotic Head with Multi-Modal Communication Capability.

AI-enabled intelligent virtual humanoids are critical to ensure natural and socially aware human-machine interactions in areas such as education, customer services, and companionship. In this paper, we develop a vision-enabled virtual robotic head which is entirely running inside a web browser and includes modules like...

Md. Ashiqussalehin, M. Khushi, Zannatul Ferdushie et al. · 0 citations
Open access 2026

MultiModal Deep Learning Framework for Missing Children Detection using Vision-Language Mode

As report of missing children continue to rise around the world there is an urgent requirement of some smart and scalable systems which can help in timely and accurate identification. In this paper, we propose a multimodal deep learning system that combines Vision-Language Transformers like CLIP, Face Re-identification...

K. U. Maheswari, P. Dhanalakshmi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.