Detecting Hidden Chain-of-Thought in Large Language Models with Linguistic, Behavioral, and Mechanistic Indicators
These findings show stronger, less prompt-conditional CoT-like behavior in the reasoning-tuned model, consistent with but not proof of latent reasoning, thus investigates latent reasoning without relying on models'self-reported traces.
Armaanjit Singh, R. Le, Jasminepreet Kaur et al.
· 0 citations