A customer-facing AI powered triaging agent that leverages large language models to conduct multi-turn conversations, ask relevant questions, and classify cases for accurate, policy-guided routing, making it embedded in the customer journey is developed.
Alankar Atreya, Stefan Sylvius Wanger, Devesh Batra et al.· arXiv.org· 0 citations
Large-scale text annotation brings expert insight to millions of documents, often through a codebook that AI annotators follow. Developing a robust codebook, however, takes months. Large language models (LLMs) could speed this process by applying an early codebook to the data, surfacing cases with strong LLM disagreeme...
Ze-Yu He, Zhu-Qian Zhou, Kirk P. Vanacore et al.· 0 citations
Safety Nudges is introduced, a browser-based tool that provides lightweight, in situ flags when concerning behavior is detected in chatbot conversations, suggesting that user facing safety nudges can complement model-level safeguards by helping people critically evaluate AI responses in context.
Varshini Elangovan, J. Wedgwood, Chhavi Yadav et al.· 0 citations
BizChat was extended, an AI-powered business-planning tool, with an evaluation module that links each generated claim to the entrepreneur's original input, and interface scaffolds like claim-to-input links primed attendees with concrete, personal evaluations.
Qi Zhao, Marjory Pineda, Ketul Chhaya et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Visual data analysis involves both open-ended exploration and targeted question answering. Visualization authoring tools support this process by enabling users to create visualizations for these tasks. With the rise of large language models (LLMs), substantial effort has been devoted to developing visualization authori...
Yuki Ueno, Bretho Danzy III, Zhuojun Jiang et al.· 0 citations
This study empirically investigates how practitioners prioritize explainability relative to four competing factors: accuracy, compliance, cost, and speed, and reveals that these priorities are structured not as a simple trade-off, but as a system of distinct prerequisites and constraints.
Patricia Marcella Evite, E. Svetlova, Doina Bucur· arXiv.org· 2 citations
Effective communication of robot touch intent is essential for safe and predictable physical human-robot interaction. While intent communication has been widely studied, existing approaches lack the spatial specificity and semantic depth necessary to efficiently convey robot touch intent. We present Mirror Skin, a ceph...
David Wagmann, Matti Kr\"uger, Chao Wang et al.· 0 citations
PRIMMDebug consists of an online tool that takes students through the steps of a pedagogical process based on PRIMM, a framework for teaching programming, and encourages written articulation throughout the debugging process and limits students'ability to run and edit code at certain stages.
Eligibility criteria play a critical role in clinical trials by determining the target patient population, which significantly influences the outcomes of medical interventions. However, current approaches for designing eligibility criteria have limitations to support interactive exploration of the large space of eligib...
Rui Sheng, Xingbo Wang, Jiachen Wang et al.· 0 citations
This study introduces an interdisciplinary framework for benchmarking robots deployed in public environments, addressing the gap between traditional laboratory metrics and real-world benchmarking requirements. We evaluate three distinct robots across diverse use cases - outdoor park cleaning, pedestrian underpass clean...
Raphael Memmesheimer, Martina Overbeck, Dominik Beyer et al.· 0 citations
Industry 5.0 (I5.0) repositions manufacturing around a human-centric vision in which cyber-physical systems must adapt to the worker rather than the other way around. Despite significant advances in worker state monitoring technologies, including markerless computer vision, wearable physiological sensors, and machine-l...
Lara Pereira, Luís Miguel D. F. Ferreira, J. Paulo· 0 citations
This work argues for instructional governance by design: governance should be encoded in a teaching tool's interaction model, constraints, and workflow, and introduces a multidimensional framework that characterizes AI teaching tools through pedagogical grounding, AI instructional authority, and human accountability an...
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduSep 30, 2026
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.