Prefilling the Reasoning Channel: Output-Prefix Attacks on Reasoning LLMs
This is the first systematic, controlled study that isolates the scratchpad reasoning channel as an output-prefix attack vector, and the first to compare reasoning-only, output-prefix-only and reasoning-plus-output-prefix attacks across both exposed- and hidden-reasoning models.
Lukáš Brůna, Robert A. Bridges, Adam Ek
· 0 citations