#artificial intelligence
May 2026
Drop the Act: Probe-Filtered RL for Faithful Chain-of-Thought Reasoning
ProFIL (**Pro**be-**Filtered Reinforcement Learning) is introduced to reduce theater, increase chain-of-thought faithfulness, and shrink chain length in a single, drop-in extension to Group Relative Policy Optimization (GRPO).
Swapnil Parekh, Naman Goyal
· arXiv.org · 2 citations