This work introduces CyberWorld, a Dreamer-style world modeling framework that learns latent cyber dynamics from vector, graph, textual, and multimodal representations of the defended network, and identifies world representation as a central design axis for robustness and scalability.
Abstract
Deep reinforcement learning has become a prominent approach to autonomous cyber defense. Existing methods are predominantly model-free and consequently require extensive environment interaction. World models provide an alternative by learning predictive dynamics and optimizing policies through imagined trajectories, yielding substantial gains in sample efficiency in robotics and embodied control. Extending this paradigm to cybersecurity raises a fundamental question: what should constitute the"world"in a cyber world model? We introduce CyberWorld, a Dreamer-style world modeling framework that learns latent cyber dynamics from vector, graph, textual, and multimodal representations of the defended network. Across all four scoreable CyberWheel attack strategies, the graph-based CyberWorld variant exceeds a strategy-agnostic control after 3.6k-15.8k environment steps, compared with millions of steps required by model-free PPO. Across representation choices, graph structure provides greater robustness under topology-dependent attacks, while simpler representations remain competitive in overall performance. Among successful runs, the number of episodes required to reach the control remains approximately constant as network size increases from 15 to 100 hosts. These results establish learned cyber dynamics as a sample-efficient and scalable basis for autonomous defense, and identify world representation as a central design axis for robustness and scalability.
An autonomous cyber defender trained with reinforcement learning (RL) is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it does not eliminate th...
Harshith Doppalapudi, Nathaniel D. Bastian, Ankit Shah· 0 citations
YAML is employed as a structured representation format for simulating complex network configurations, thereby enabling Large Language Model-driven pipelines to support and improve reinforcement learning (RL) agent training, and underscore the transformative potential of integrating LLMs into cybersecurity research.
S. Kampakis, Fabio Rovai, Marcos Charalambides et al.· 1 citation
To achieve effective, stealthy, and persistent control, TrojanWorld combines Decision-Reflective Induction to steer trigger-conditioned imagination toward attacker-specified actions using decision feedback, Clean Behavior Anchoring to preserve trigger-free predictive and behavioral fidelity, and Causal Propagation to s...
Wen-Kai Huang, Si-Yuan Liang, Gaolei Li et al.· 0 citations
As cyber threats to power grid infrastructures escalate, the urgency of understanding how to protect cyber-physical systems (CPS) has never been greater. These systems, which integrate physical processes with digital control, are increasingly susceptible to sophisticated cyberattacks that can lead to widespread disrupt...
A. Raptis, S. Gritzalis, A. Yannacopoulos· International Journal of Inf...· 0 citations
Training capable cyber agents is often treated primarily as a problem of model scale, yet open-weight post-training is constrained more directly by the cost of executable environments, reliable multi-turn supervision, and access to strong teachers. We present a data-centric framework that addresses these bottlenecks th...
Zong-Jie Li, W. AlanZ, J. JohnNicolas et al.· 0 citations
BlueSTAR is presented, a tiered agentic architecture for autonomous cyber defense in enterprise IT/OT networks that retains the fast containment of deterministic response for known threats while successfully defending against attacks requiring contextual and cross-cycle reasoning.
Simona Boboila, Xavier F. Cadet, Edward Koh et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 29, 2026
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.