Cybersecurity Detection Classification with Reasoning-enabled Language Models
This work trains a chain-of-thought reasoning-enabled triage classifier on real, human-labeled Windows endpoint detections by combining automated prompt optimization, self-training, and reinforcement learning with verifiable rewards, and shows that a finetuned 30B model significantly outperforms frontier general-purpose models, motivating targeted training over scale.