HarnessSecurity-Bench: Do Security Mechanisms Really Protect Coding Agent Harnesses?
This work presents HarnessSecurity, the first systematic empirical study and benchmark of open- and closed-source coding agent harnesses, and derives a ten-mechanism taxonomy and assesses 400 harness-mechanism cells using independent ratings by researchers and large language model judges.