FairAware, a fairness assessment tool co-designed with Human Resources domain experts, is presented, suggesting that fairness assessment tools for non-experts are usable for identifying biases but need built-in checks on understanding before stakeholders make higher-stakes decisions.
Anna Verheyden, Yi-Zhe Zhang, Robin De Croon et al.· 0 citations
Given recent achievements of large language models (LLMs), frontier models are expected to perform well on Bayesian reasoning tasks, at least as well as humans. Furthermore, there is no reason to expect that LLMs will condemn others who offer those very same Bayesian judgments, a fallibility observed in human decision-...
Human intention can be modeled as a latent internal state that modulates how sensory information is evaluated and translated into action in human-machine systems. However, most existing brain-computer interfaces (BCIs) rely on control signals tightly coupled to externally imposed stimulation and do not explicitly infer...
Xiaowei Jiang, Daniel Leong, Yu-Cheng Chang et al.· 0 citations
Wearable AI/ML research needs raw, synchronized, and reconfigurable multimodal data, but consumer devices are closed and many research platforms remain tied to one embodiment or sensor set. This paper presents SensWear, an open, modular, and AI-ready wearable platform that decouples embodiment, sensing, data interfaces...
Bullying in schools profoundly affects the mental and physical health of teenagers. Although existing in-person and digital interventions provide some benefits, they often fall short in addressing the complex social dynamics of bullying. In this study, we collaborated with K-12 teachers to co-design a multi-agent anti-...
Jiaju Lin, Ellen Wenting Zou, Feiwen Xiao et al.· 0 citations
The effectiveness of edge-cloud collaboration for GUI grounding depends on autonomous requesting, where the edge agent selectively offloads complex tasks to the powerful cloud. However, in visually dense scenarios, lightweight edge agents often exhibit overconfident hallucinations, leading to a misalignment between con...
This work introduces a temporal-graph-learning (TGL) controller that learns these predictions jointly and supplies both outputs in one forward pass, and shows that a small graph model can handle both decisions and improve the language agents it controls.
Xiaoze Liu, Ruowang Zhang, Amir H. Abdi et al.· 2 citations
This work provides a general theoretical characterization of the problem of learning to assign prediction tasks to one agent from a set of available agents, including human decision-makers and AI models and develops a framework of sequential explore-exploit policy-learning algorithms that seek to maximize overall perfo...
People increasingly turn to general-purpose AI chatbots for advice about emotional and mental health problems, but the ability of these systems to recognize and appropriately triage psychiatric emergencies remains under-characterized.
We evaluated psychiatric triage performance in 15 frontier AI chatbots using 112 cl...
Veith Weilnhammer, Lennart Luettgau, Christopher Summerfield et al.· 0 citations
We present the first motion generation system for playtesting virtual reality (VR) games. Our player model generates VR headset and handheld controller movements from in-game object arrangements, guided by style exemplars and aligned to maximize simulated gameplay score. We train on the large BOXRR-23 dataset and apply...
Nam Hee Kim, Jingjing May Liu, Jaakko Lehtinen et al.· 0 citations
Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We a...
The rapid, unpredictable advancements in AI system capabilities has seen regulators take adaptive and experimental approaches to policymaking. Established in other domains as instruments balancing regulation with innovation, regulatory sandboxes are seen as solutions for AI regulation. However, analyses mostly focus on...
Idoia Landa-Oregi, Tom Deckenbrunnen, Alessio Buscemi et al.· 0 citations
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduSep 30, 2026
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.