Evaluating Coding Agents on Kernel Exploit Generation
This work introduces KEX-bench, a benchmark for evaluating coding agents on exploit primitive generation against real operating-system kernels, and evaluates state-of-the-art coding agents paired with frontier and open-weight models under fixed tool-call budgets.
Junyoung Jang, Gwanhyun Lee, Hwiwon Lee et al.
· 0 citations