Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents
The red-teaming methodology and highlighting new attack vectors aim to help defenders evaluate their mitigations against the possibility of persistent malign coding agents and prevent multi-context attacks at an acceptable cost.