How MIT students are helping to prevent cyberattacks
Students from the MIT Cybersecurity Clinic help local governments and other vulnerable organizations defend against digital threats.
More from the blog
3 Questions: What is the best path forward for AI in academia?
MIT Statistics and Data Science Center Director Alexander (Sasha) Rakhlin shares important considerations for departments and institutions.
Using AI to mitigate the growing environmental threat of data centers
By rethinking how large cloud computing systems operate, Associate Professor Christina Delimitrou seeks to make data centers more energy efficient.
Discovering the value of humanistic inquiry
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Chris Bourg named vice provost and Barbara K. Ostrom (1978) Director of the MIT Libraries
As director, Bourg has focused on digital access, open and equitable scholarly publishing, and expanded support for data-intensive research.
Related papers
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
A novel threat is unveiled in which attackers steer the RAG system's response by injecting malicious passages into its knowledge base, enabling the attacker to steer the response without altering the user input or modifying the RAG weights.
OverThink: Slowdown Attacks on Reasoning LLMs
This work evaluates Overthink on proprietary and open-source reasoning models across the FreshQA, SQuAD, and MuSR datasets, and shows that newer generations of RLMs, while showing a drastic increase in per-token cost, also exhibit up to a 2.3x increase in reasoning tokens, leaving them more vulnerable to Overthink atta...
Learning diverse attacks on large language models for robust red-teaming and safety tuning
This work proposes to use GFlowNet fine-tuning followed by a secondary smoothing phase, to train the attacker model to generate diverse and effective attack prompts, and finds that the attacks generated by the method are effective against a wide range of target LLMs, both with and without safety tuning, and transfer we...
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
The key findings show that while some detectors can identify attacks that rely on explicit textual instructions or visible image perturbations with moderate to high accuracy, they largely fail against attacks that omit explicit instructions or employ imperceptible perturbations.