From Capability to Assurance in Autonomous Penetration-Testing Harnesses: A Framework and Reference Implementation
Research on large language model agents for penetration testing is evaluated almost entirely by capability: whether the agent captures a flag or reproduces a proof of concept. That metric suits a benchmark but is silent on the properties that decide whether an autonomous agent can be used in an authorized engagement: w...