PhysDox is introduced, a physical feasibility auditing benchmark for biomedical protocols comprising a 683-sample expert-curated Gold set and a 5,000-sample Silver set across six sensing domains, demonstrating that protocol auditing demands calibrated feasibility reasoning rather than factual recall or longer rationale...
He Liu, Boyuan Gu, Shuai Cheng et al.· arXiv.org· 2 citations
Developing a novel research idea is hard. It must be distinct enough from prior work to claim a contribution while also building on it. This requires iteratively reviewing literature and refining an idea based on what a researcher reads; yet when an idea changes, the literature that matters often changes with it. Most...
Hita Kambhamettu, Bhavana Dalvi Mishra, Andrew Head et al.· 0 citations
AI-powered road surveillance systems are increasingly proposed to monitor infractions such as speeding, phone use, and jaywalking. While these systems promise to enhance safety by discouraging dangerous behaviors, they also raise concerns about privacy, fairness, and potential misuse of personal data. Yet empirical res...
Ziming Wang, Shiwei Yang, Rebecca Currano et al.· 0 citations
Post-traumatic stress disorder (PTSD) is associated with sudden, uncontrollable, and intense flashbacks of traumatic memories. Trauma exposure psychotherapy has proven effective in reducing the severity of trauma-related symptoms. It involves controlled recall of traumatic memories to train coping mechanisms for flashb...
Annalisa Degenhard, Stefan Tsch\"oke, Michael Rietzler et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Talk2Escape is introduced, a proactive and model-agnostic dialogue intervention framework that reframes navigation as a closed-loop interactive process and proves that proactive dialogue drastically improves navigation robustness in physical environments.
Ze-Rui Li, Si-Hao Lin, Yan-Yan Shao et al.· 0 citations
An LLM-assisted workflow that represents baseline and current requirements as triple-based semantic graphs and supports side-by-side comparison of curated graph snapshots and uses visual encoding to highlight structural changes.
K. McFarland, Song-Hui Yue· IEEE International Requireme...· 0 citations
Electrovibration-based tactile feedback is demonstrated to be a viable and effective modality for robot teleoperation, improving operator responsiveness and sense of presence in contact-rich manipulation tasks, with direct applicability to safety-critical domains such as nuclear maintenance.
Alperen Kenan, Juan Jose Garcia Cardenas, Adriana Tapus et al.· 0 citations
Avatar-streaming systems are commonly evaluated with image and video quality assessment (IQA/VQA) metrics, implicitly treating visual fidelity as a proxy for communicative success. We test this assumption through a controlled behavioral study of rendered 3D avatars across a pristine condition and fourteen geometric, ph...
This work systematically analyzes all 1,702 CHI 2026 full papers and identifies 125 that annotate videos, and derives a five-dimensional taxonomy spanning analytic purpose, viewpoint, phenomenon, reasoning requirement, and annotation authority to map the capabilities and limitations of a general-purpose VLM.
Xi-Yuan Shen, Jiuyang Lyu, Seokhyun Hwang et al.· 0 citations
The objective of this article is to provide design principles and a software architecture for enabling interaction between humans and multiple agents in simulated dynamic worlds. This connects the current era of general artificial intelligence (AI/AGI) with the proliferation of transformer-based conversational agents a...
Human-AI Teaming (HAT) reviews often group studies by labels such as advisor, teammate, or coordinator. Yet the same label can describe one person taking AI advice, several people coordinating around AI, or a workflow that distributes authority and responsibility. Pooling these studies can therefore change the human un...
Hanjing Shi, Kimberly Wang, Sabrina Doherty et al.· 0 citations
As AI systems are increasingly integrated into professional work, reflection strategies such as cognitive forcing and prompts that foster critical engagement have shown promise in reducing overreliance and improving decision quality. However, these strategies have primarily been evaluated as short-term interventions wi...
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduSep 30, 2026
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.