Jul 2026· International Conference on Future Internet of Things and Cloud· pp. 441-448· 0 citations· 11 references
Abstract
This paper proposes a well-rounded media consumption system that utilizes speech recognition and generative Artificial Intelligence to create a user-influenced slideshow. The system runs on a Raspberry Pi, using a microphone for input and a touchscreen for output. Spoken user requests are transcribed and categorized into one of five content classes: news, daily activities, social media trends, interest-based topics, and storytelling. These categories are then converted into prompts suitable for image generation. The system uses AI to condense unstructured speech into descriptive, image-ready content, reducing cognitive overhead and minimizing screen time. Unlike conventional browsing, this approach enables passive, voice-controlled consumption of highly relevant media. Qualitative and quantitative evaluation shows that the system reliably transcribes varied speech inputs, classifies user intent with high interpretability, and generates coherent, category-aligned visuals. Specifically, the LLM-based intent classifier achieved 90% accuracy and a Macro-F1 of 0.931 on the tested prompts, while on-device transcription operated at an average CPU utilization of 7.3% with a peak SoC temperature of 48.7°C, confirming the feasibility of the pipeline on a Raspberry Pi. The use of AI in this context enhances personalization, reduces interaction friction, and supports timeefficient engagement with digital media.
An automated pipeline for transcription and summarization of video conferences held on the Jitsi Meet platform that requires no proprietary cloud API keys and is deployable on-premise via Docker Compose, making it suitable for organizations with strict data-privacy requirements.
G. Amirkhanova, L. Bektemir, Shyrailym Adilkyzy et al.· International Conference on...· 0 citations
Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatars, and embodied agents. While recent audio-video generative models can synthesizehigh-fidelity synchronized content, existing supervision is largely \emph{descriptive}:models are trained to render captions rather than to produce audio-visual responsescaused by external user interactions. We introduce \textbf{InteracVid}, \emph{the firstopen-source large-scale dataset that addresses this missing supervision}, so that everysample couples a preceding audio-visual context and an external stimulus with the realinteractive response that follows. We design a metadata-aware pipeline that extractsinteractive clips from long, noisy livestreams, yielding over \textbf{454K}context-query-response triplets from more than \textbf{59K} livestream videos andspanning conversation-centered, object-centric, procedural, embodied, and screen-basedscenarios. A ten-rater human study confirms that the extracted interactions are causal,natural, and temporally complete for both genuine and reconstructed queries. On aheld-out benchmark of \textbf{100} genuine live-chat queries, fine-tuning on InteracVidimproves both interaction planning and audio-video response generation, and anindependent human evaluation reproduces the system ranking and the conclusions obtainedwith our automatic judge. These results highlight interaction-structured data as acritical foundation for interactive multimodal generation.
Chi Zhang, Hao-Yan Shi, Yueyi Liu et al.· 0 citations
People with physical and motor disabilities face significant barriers when interacting with image creation applications that rely on keyboard input, touchscreens, or complex graphical interfaces. Although modern AI image generation models produce high-fidelity images from textual descriptions, they remain dependent on typed input, which is inaccessible to users with limited hand mobility. To address this gap, this paper proposes a voice-to-image assistive system that enables users to create images exclusively through voice commands. The system integrates three tightly coupled modules: (i) a Speech-to-Text (STT) engine based on OpenAI Whisper, which transcribes spoken commands into raw text; (ii) an NLP-based prompt refinement module that applies tokenization, grammatical correction, and semantic augmentation to produce diffusion-friendly image prompts; and (iii) a Stable Diffusion image synthesis backend that generates high-resolution images from the refined prompts using GPU-accelerated latent diffusion. The system achieves STT accuracy of up to 100% for simple commands and 96% for long descriptive inputs, with prompt-to-image semantic alignment reaching up to 97%. Average image generation time is 6–12 seconds on GPU hardware. This work demonstrates the practical viability of multimodal AI pipelines for assistive applications and outlines directions for future development, including multilingual support, mobile deployment, and integration with existing assistive technology ecosystems.
Dasari Sri Krishna, Vaddi Radhesyam, Lakshmi Aiswarya Narikimilli et al.· 2026 4th International Confe...· 0 citations
Taking notes during meetings sounds easy, but in reality, people often miss key points—especially when multiple people are talking. Managing follow-ups in separate apps only adds to the hassle. MinuteMaster was built to simplify this. It's a web app that brings transcription, speaker identification, summarization, and scheduling into one place. It uses Whisper for multilingual speech-to-text, pyannote.audio to identify who's speaking, and BART to turn long transcripts into clear, short summaries. A built-in calendar helps users manage meetings without switching apps. We tested it on 30 recordings across English, Hindi, Telugu, Tamil, and mixed languages. Transcription accuracy ranged from 76\% to 88\%, and summaries performed better than TextRank with a ROUGE-1 score of 0.48. In a study with 42 users, it received an average rating of 4.25/5. The system runs on Flask with MongoDB, making it simple to deploy and scale.
B. Banik, Pampari Yukitha, Charan Sai Venna et al.· Next-Generation Computing Sy...· 0 citations
A user study comparing two modalities for writing prompts for generative AI tasks reveals that input modality significantly influenced prompting behaviour but did not lead to measurable differences in subjective evaluations.
Nishant Rathore, Tushar Billakanti, J. Ceha et al.· International Conference on...· 0 citations
The generative artificial intelligence has made tremendous advancement in visual content generation; but the issue of intuitive and user-friendly adjustment of the current video scenes is a difficult question to answer. The paper describes a personalized movie reimagining system based on AI that reimagines input video scenes based on natural language instructions. In contrast to the traditional text-to-video methods, which make the content anew, the suggested structure of the visual content is based on structure-preserving transformation, meaning that users have the opportunity to adjust visual properties while preserving the original composition of the scene and motion dynamics. It combines computer vision to comprehend the scene, natural language processing to decipher user intent, and diffusion-based generative models to synthesize videos. ControlNet using Canny edge conditioning is used to maintain spatial structure and motion adapters in AnimateDiff maintain temporal coherence across frames. Also, a proactive enhancement mechanism and dynamic conditioning plan enhance congruence between user input and output generated. The results of the experimental work conducted on a variety of scene types indicate that the proposed system reaches a semantic alignment score of 4.2/5 and structural similarity index (SSIM) of 0.81, which is better than the baseline text-to-video-based methods. The system produces the short video sequences (8-24 frames) in 120-180 seconds with the standard GPU hardware. These results indicate the usefulness of the framework in facilitating high-quality video transformation.
Thupakula Venkata Sumanth, Kriti Gupta, Charanjit Singh et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.