Automated PDF accessibility checkers provide useful conformance evidence, but a clean report is not a complete accessibility oracle. PDFa11yMut measures mutation-specific checker behavior by applying paired structure-level transformations to reference-suite baselines, verifying intended deltas and non-target invariants...
Reliability in medical Q&A is often pursued by grounding responses in authoritative medical information. We show that when Q&A is embedded within ongoing care, reliability depends on more than what the system knows medically. In a study with 73 skin cancer patients practicing postoperative wound care, 41.9% of response...
The effects of scent-delivery within virtual reality (VR) environments by using fan-mediated olfactory devices are explored, indicating that scent-delivery feedback influenced spatial behavior in distinct ways across the two studies.
Siyeon Bak, Junho Kim, Dong-Yun Han et al.· 0 citations
VocalEyes is built, a speaker-aware AR captioning system that creates named voice profiles from natural self-introductions that shows how in-conversation registration can preserve speaker attribution across live captions and meeting records in scripted small-group meetings.
Yu-Xiao Wang, Xu-Long Tang, Chen Chen et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Results show that olfactory cues can support recognition of encountered locations during navigation, while ventilation was associated with reduced residual interference and improved usability, however, longer ventilation did not clearly improve recognition accuracy in repeated navigation.
Dong-Yun Han, Heecheol Kim, Siyeon Bak et al.· 0 citations
A preregistered, three-country randomized experiment on out-of-context image misinformation, manipulating a correction's channel affordances (synchronicity, bandwidth) across four conditions shows engagement mechanisms do not substitute for slow AI literacy.
Kokil Jaidka, I. Mujtahid, Peng Qi et al.· 0 citations
Environmental interfaces are introduced: designed environmental conditions that shape human-AI relationships through ambient, holistic, and evaluative pathways rather than explicit functional interaction, pointing to a broader transition from designing interfaces for operating intelligent systems to designing environme...
Ke-Qi Chen, Run-Jia Tan, Xin-Yi Fu et al.· 0 citations
The first user-centric benchmark grounded in real user feedback signals for evaluating preference alignment and dialogue generation is presented, demonstrating that user feedback prediction is a learnable capability and revealing different aspects that influence user experience.
Mengze Hong, Xia Zeng, Zeyang Lei et al.· arXiv.org· 1 citation
Analysis of reasoning errors in vision language models and how they impact user trust and the ability to detect errors highlights how CoT explanations can simultaneously clarify and mislead, underscoring the need for NLP systems to provide explanations that encourage scrutiny and critical thinking rather than blind tru...
Eunkyu Park, Wesley Hanwen Deng, Vasudha Varadarajan et al.· arXiv.org· 6 citations
This work introduces NSV-Shift, a contrastive benchmark for evaluating whether speech-to-speech models can understand non-speech vocalizations (NSVs) and adapt their responses accordingly, and evaluates five models on NSV perception, emotion understanding, and response adaptation.
The results suggest that growth helps when its criterion can rank the candidate neurons, and that the rate of skipped neuron addition tells where a decoder can be grown small from scratch.
Adam Mounir, Stella Douka, A. Caillet et al.· 0 citations
Through a four-speaker attention decoding benchmark, it is shown that combining behavioral and physiological signals improves decoding performance over EEG-only approaches, enabling future advances in multimodal auditory attention decoding.
K. M. Naimul Hassan, Ali Alavi, D. Williamson· 0 citations
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduSep 30, 2026
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.