Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents
The red-teaming methodology and highlighting new attack vectors aim to help defenders evaluate their mitigations against the possibility of persistent malign coding agents and prevent multi-context attacks at an acceptable cost.
Alex Remedios, Simon Storf, Fabien Roger et al.
· 0 citations