Self-play red-teaming improves language-model safety by pitting attacker and defender roles against each other in a zero-sum game. However, real adversaries increasingly use mosaic attacks: multi-turn sequences whose individual fragments are innocuous in isolation yet assemble into a harmful payload. We develop a theory of mosaic defense that characterizes what is required to prevent such attacks without sacrificing helpfulness. We first show that no fixed bounded window of recent prompts is sufficient in general: safety-relevant information may occur arbitrarily far back in the interaction. We formalize a watchman, an online state mechanism that carries this information forward, and show that under explicit assumptions it enables zero-failure defense with positive benign helpfulness. Under stronger conditions, it is also optimal among zero-failure defenders. An exact watchman may nevertheless require exponentially many states, while exact maliciousness detection can require exponentially many queries in an unstructured black-box model. These state and query lower bounds do not by themselves imply hard learning: the construction underlying the state lower bound is efficiently learnable from labeled examples, whereas certifying worst-case safety can require substantially more information under restricted access. We also show that self-play equilibrium alone does not certify usefulness, motivating a constrained formulation that maximizes worst-case benign helpfulness among zero-failure defenders. Empirically, training role-specific attacker and defender LoRA adapters over frozen LLMs via multi-turn self-play strengthens both roles: attackers become more effective at eliciting harmful responses, while defenders become more robust to attack, with improvements also observed on unseen attack objectives.
The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.
Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al.· Information and Software Tec...· 394 citations· ⚡54
The possibility of inferring high-dimensional data inference in a model that consists of a prior and an auxiliary differentiable constraint given some additional information is considered, thereby allowing a range of potential applications in adapting models to new domains and tasks.
Alexandros Graikos, Esmeralda S. Whitammer, N. Jojic et al.· Neural Information Processin...· 316 citations· ⚡15
It is proved that any global minimizer of the trajectory balance objective can define a policy that samples exactly from the target distribution, and empirically demonstrate the benefits of the trajectories balance objective for GFlowNet convergence, diversity of generated samples, and robustness to long action sequenc...
Esmeralda S. Whitammer, Moksh Jain, Emmanuel Bengio et al.· Neural Information Processin...· 302 citations· ⚡60
GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations...
Xiaotian Zhang, Chun-yan Li, Yi Zong et al.· arXiv.org· 216 citations· ⚡17
This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.
Carmine Giardino, Xiaofeng Wang, P. Abrahamsson· International Conference on...· 175 citations· ⚡19
This work investigates the possibilities of using LLMs in a resume screening setting via a document retrieval framework that simulates job candidate selection and finds that the MTEs are biased, significantly favoring White-associated names in 85% of cases and female-associated names in only 11.1% of cases.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.