The integration of Artificial Intelligence (AI)/Machine Learning (ML) into high-impact domains such as finance and autonomous systems offers significant benefits, but also introduces complex risks and regulatory challenges. These systems exhibit properties including non-determinism, data dependence, and evolving vuln...
Anita Khadka, C. Maple· Artificial Intelligence Revi...· 0 citations
Large language models (LLMs) are capable of completing a variety of tasks, but remain unpredictable and intractable. Representation Control (RepControl) seeks to resolve this problem through targeted interventions that modify high-level representations of concepts such as honesty, harmfulness or power-seeking. We forma...
LLM agents increasingly take privileged, often irreversible structured actions, such as paying an invoice. They assemble each action from action-critical fields in documents and tool outputs that an adversary can corrupt, and indirect prompt injection can drive the model itself to extract attacker-chosen values. Curren...
Anmol Pandey, A. Jain, Liang-Wei Chen et al.· 0 citations
Deep reinforcement learning (DRL) enables adaptive intrusion detection in dynamic network environments but also exposes intrusion detection systems (IDS) to adversarial threats such as universal adversarial perturbations (UAPs), which apply a single input-agnostic perturbation to degrade detection performance across tr...
Hong-Sen Zhang, Lu Zhang, Ming-Jing Xu et al.· 0 citations
Ethereum decentralized finance (DeFi) provides a public, time-stamped record of transaction-level event streams, but the same public symbols can create strong machine-learning shortcuts, so ETH-TraceBench treats difficult transfer and controlled-input conditions, rather than a single aggregate score, as the main evalua...
These results show that conversational refusal benchmarks can substantially overstate the safety of deployed coding agents and motivate defenses that reason about safety across multi-turn IDE workflows and their generated artifacts, not only individual chat turns.
This work proposes PRoVeFL-a novel, modular FL framework that is Privacy-preserving, Byzantine-Robust, and ensures Verifiable aggregation, and improves runtime over the prior works, Prio and ELSA, based on distributed trust with comparable security guarantees, up to 100x and 10x, respectively.
Harsh Kasyap, Anil Kumar Pradhan, U. Atmaca et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.