Learner engagement is commonly viewed as a key factor in successful learning. In online settings, limited face-to-face interaction can make learners more prone to distraction and reduced attention, highlighting the importance of monitoring and sustaining engagement. Recent advances in generative AI allow systems to infer learners’ cognitive and emotional states from multimodal cues, enabling more personalized and adaptive instructional support. However, little research has examined how such systems can dynamically adapt both the type and timing of feedback based on learners’ moment-to-moment engagement states inferred from multimodal signals. This work presents an adaptive multimodal AI-driven tutoring system that infers learners’ states by interpreting real-time visual, auditory, and behavioral cues. Based on the inferred learner state, the AI tutor determines when and how to intervene to sustain engagement. The system is structured as a closed-loop cognitive architecture: perception (capturing real-time multimodal cues), decision (aggregating the multimodal inputs into four affective metrics), and action (delivering feedback based on the inferred state by mapping each metric to a feedback type and timing strategy). This work presents a high-fidelity, interactive multimodal AI tutoring system that illustrates the feasibility of integrating multimodal cues to enable adaptive instructional feedback and engagement-aware intervention in online learning contexts.
Jaewon Jung, Aahil Shaikh, J. Shaughnessy et al.· Proceedings of the Human Fac...· 0 citations
Visual impairments affect upwards of 2.2 billion people worldwide. As AI systems increasingly support navigation for people with visual impairments, how uncertainty is communicated becomes critical. Prior work shows that communicating uncertainty can improve trust calibration and decision-making, yet it remains underexplored in assistive navigation. This project develops an uncertainty-aware assistive navigation architecture that integrates AI uncertainty into auditory guidance during real-time scene descriptions. Rather than using explicit confidence statements, the prototype embeds uncertainty into speech via variations in tone, pacing, and emphasis. The prototype combines a real-time collision-warning module with a semantic reasoning layer powered by a large language model (LLM). When generating scene descriptions, token-level uncertainty is mapped to auditory prosodic cues, enabling users to implicitly gauge the system’s confidence without disrupting navigational task flow. This work presents a high-fidelity prototype that treats AI confidence as an interaction design feature, illustrating how model uncertainty can be rendered perceptible in assistive navigation and reframed as a human-factors design parameter.
Hayden Shaffer, Aahil Shaikh, He Zhang et al.· Proceedings of the Human Fac...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.