Natural Language Understanding (NLU) plays a crucial role in various applications, yet its performance suffers from weaknesses in handling the complexities of human languages, ranging from lexical ambiguity to high-level reasoning difficulties. Analyzing errors across diverse linguistic phenomena is crucial for NLU imp...
Khloud Al Jallad, Nada Ghneim, Ghaida Rebdawi· 0 citations
Data videos communicate data insights through dynamic charts, voice narration, and synchronized animations, and have become a widely adopted form of data storytelling. However, producing them requires expertise in data analysis, narrative design, and video editing. Static visualization tools lack narrative and animatio...
Yu-Peng Xie, Zhen-Yang Wang, Liang-Wei Wang et al.· 2 citations
This work argues that their continual coordination under competing demands constitutes an important and underexplored target for modern robot learning and proposes the ethological behavioral substrate as a conceptual lens for studying this form of competence in artificial agents.
As coding assistants become increasingly autonomous, developers run multiple sessions in parallel, shifting the challenge from code generation alone to coordinating and monitoring concurrent agent work. Through a formative study (N=14), we identified PILOT: five supervisory practices for Planning, Isolating, Logging, O...
Tao Long, Wei Shi, Hussein Mozannar et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Interactive dashboards require users to reveal and connect evidence across stateful interactions. Although graphical user interface (GUI) agents could automate this process, existing dashboard benchmarks primarily report final answers or task success. They provide limited insight into whether failures arise from mainta...
Chu-Han Zhang, Qinghongbing Xie, Zi-Yue Wang et al.· 0 citations
These findings show that Vis designers actively use LLMs for both creative and technical aspects of the visualization process, and opens up opportunities for research combining LLM-mediated work with Vis tools that incorporate data visualization guidance, constraints, and best practices.
S. C. Spivak, Aditi Krishna, Mahsan Nourani et al.· IEEE Computer Graphics and A...· 0 citations
PainterBench, a benchmark that ports the incomplete-drawing task to the agentic setting, is introduced and ViDrA-adapted, an automated scorer that predicts human creativity ratings of agent drawings, is presented.
Shane K. A. Dalumura Hettige, J. Oppenlaender· 0 citations
Healthcare administrative staff transfer structured information from electronic health records, referrals, claims systems, provider rosters, and work queues into dynamic forms. We developed and evaluated CLAIRE (Clinical Language and Agentic Intelligence for Reasoning and Entry), a hybrid workflow that separates field-...
We investigate the effectiveness of artificial intelligences (AI)-specifically large language models (LLMs)-relative to human scientists at high-level cognitive tasks in social science such as theory formulation, predictions of novel empirical results, and theory revision in response to new evidence. The research domai...
Ke Li, Spyros I. Zoumpoulis, Phanish Puranam et al.· 0 citations
Large language models (LLMs) are increasingly integrated into educational settings, yet educators lack robust, standards-aligned tools to evaluate their effectiveness in K-12 science contexts. Existing benchmarks predominantly assess general language or advanced scientific reasoning, leaving a critical gap in understan...
Noah L. Schroeder, Yessy Eka Ambarwati, Yu-Ji Zhang et al.· 0 citations
Teleoperated ultrasound can improve diagnostic medical imaging access for remote communities. Having accurate force feedback is important for enabling sonographers to apply the appropriate probe contact force to optimize ultrasound image quality. However, large time delays in communication make direct force feedback im...
Ryan S. Yeung, David G. Black, Septimiu E. Salcudean· 0 citations
It is demonstrated that the fixed reality modality affects HRI results, and dynamically changing the modality along the RVC improves them, highlighting the value of adaptive XR interfaces for human-robot symbiosis.
Carl Tornberg, Alicia Torck, Lotfi El Hafi et al.· 0 citations
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduSep 30, 2026
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.