This work formalizes a different design target, structured unpredictability, as conditional dependence between an output and a persistent hidden state beyond what an observer can infer from the transcript, as conditional dependence between an output and a persistent hidden state beyond what an observer can infer from t...
In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evidence about grounding across two such tasks. Working from discrete behavioral annotations, we map HCRC MapTask (Anderson et al., 1991) and MUNDEX (T\"urk et al., 2023) int...
It is suggested that chatbot communication style influences users' perceptions of conversational agents and may improve performance relative to less supportive chatbot designs, but the overall value of chatbot interaction depends on the task context.
Erik Derner, Dalibor Kučera, Aditya Gulati et al.· Computers in Human Behavior· 0 citations
This survey consolidates and analyzes developments across EEG-to-image synthesis, EEG-to-text generation, and EEG-to-audio reconstruction, and highlights open-source datasets and baseline implementations to facilitate systematic benchmarking and accelerate progress in EEG-driven neural decoding.
Shreya Shukla, Jose Torres, Akshaj Murhekar et al.· arXiv.org· 10 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Computer-use agents increasingly operate software designed for people, but interfaces often leave actions or task state unclear to machine readers. We present Affora, a design system that supports both readers while preserving visual freedom and familiar human workflows. Three controlled studies examine component imple...
GYROval - Gridded Yielding of Robust value Orientation is presented, together with the results of administering it to twenty models, a robust benchmark for measuring cultural value orientation in large language models on the two Inglehart-Welzel axes over several domains and roles.
Alexander Didenko, A. Shabanova, Vladislav Zapylikhin et al.· 0 citations
Organizations increasingly use oversight loops where one large language model (LLM) audits another's outputs alongside procedural traces of claimed steps. A common concern about such LLM-as-a-judge pipelines is that detailed traces make overseers gullible. Using signal detection theory, we audit five LLM overseers on 1...
This work analyzes three public corpora: CoAuthor (1,447 keystroke-level co-writing sessions), RealHumanEval (editor telemetry from 243 programmer records), and a pre-LLM CS1 corpus as a human-only baseline, comparing minimal-AI work, collaborative AI use, and simulated wholesale delegation.
This work introduces Lexara-RF, a reference-free set of metrics that scores CVA outputs using only the prompt, data, and model response, and reformulate evaluation as verification: 13 metrics operationalize visualization design theory and Gricean cooperative principles as computable consistency, intent-alignment, and d...
AI is changing what leaders must judge, explain, learn, and coordinate, yet existing measures do not capture these behaviors at the level needed to study leadership in AI-enabled work. We develop the AI Leadership Battery, which organizes 36 behaviorally specific subdimensions into 11 theory-specified content families....
Learning from human feedback has become a central paradigm for training modern AI systems, where models of human utility are used as reward models in policy learning. Existing methods typically assume a \emph{universal utility} function shared across a population and treat disagreement between annotators as stochastic...
Shiwali Mohan, Matthew K. Hong, Du-Le Shu et al.· 0 citations
A few-shot learning approach is developed and a small subset of question types accounts for the majority of student inquiries, and that the types of questions students ask change substantially as the task progresses.
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduSep 30, 2026
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.