Human-agent evaluations often compress interaction into a single performance score, even when human and automated policies adapt differently over time. We study this in a sequential cyber-defense game on an attack graph, where a human or reinforcement-learning defender protects cloud assets against a Deep Q-Network (DQN) attacker. We compare four defender settings: a human reward-only game operationalizing the DQN information structure, a human reward-plus-transition game operationalizing the Cognitive Hierarchy Theory-driven DQN (CHT-DQN) information structure, an automated DQN defender, and an automated CHT-DQN defender. In the reward-only game, participants receive payoff and reward feedback. In the reward-plus-transition game, they also see attacker-aware transition probabilities from the CHT-DQN model. Across 80 Mechanical Turk participants and matched automated simulations, human defenders adapted to outcomes in a way the automated defenders did not: they were more likely to reselect a node after a successful defense than after a failure, in both games, an asymmetry we interpret as consistent with Prospect Theory and Cumulative Prospect Theory. The 40-round average also mixes early rounds, where the DQN attacker acts mostly at random, with late rounds, where it mostly exploits; in the final stage, mean protection ranks the reward-plus-transition human game above the reward-only human game, then the automated CHT-DQN defender, then the automated DQN defender, an ordering the overall mean hides. By contrast, the reward-plus-transition game does not produce a statistically reliable overall gain in weighted data protection over the reward-only game. These results suggest that evaluating human-agent cyber-defense systems only by an averaged task score can miss behaviorally meaningful differences in adaptation, action allocation, and bounded human reasoning.
This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.
P. Abrahamsson, O. Salo, Jussi Ronkainen et al.· arXiv.org· 727 citations· ⚡54
The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.
M. Pikkarainen, Jukka Haikara, O. Salo et al.· Empirical Software Engineeri...· 401 citations· ⚡48
The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.
Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al.· Information and Software Tec...· 394 citations· ⚡54
The possibility of inferring high-dimensional data inference in a model that consists of a prior and an auxiliary differentiable constraint given some additional information is considered, thereby allowing a range of potential applications in adapting models to new domains and tasks.
Alexandros Graikos, Esmeralda S. Whitammer, N. Jojic et al.· Neural Information Processin...· 316 citations· ⚡15
It is proved that any global minimizer of the trajectory balance objective can define a policy that samples exactly from the target distribution, and empirically demonstrate the benefits of the trajectories balance objective for GFlowNet convergence, diversity of generated samples, and robustness to long action sequenc...
Esmeralda S. Whitammer, Moksh Jain, Emmanuel Bengio et al.· Neural Information Processin...· 302 citations· ⚡60
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
MIT News · Artificial Intelligence· news.mit.eduSep 25, 2026