COMPASS: Finding Where Reasoning Lives in Language Models
Comparing three model families and multiple math benchmarks, COMPASS outperforms the activation-steering baselines the authors compare against, improves GSM8K accuracy by 16 percentage points on average, and approaches CoT accuracy with 20-70\% fewer generated tokens.
Pratyay Dutta, Kowshik Thopalli, V. Narayanaswamy
· 0 citations