Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
The results suggest that CoT monitoring may be less reliable when preference information arrives through tools or must be inferred from raw artifacts, within the single-call, prefilled-tool setting tested here.
Aryo Pradipta Gema, Neel Rajani, Rohit Saxena et al.
· 0 citations