#artificial intelligence
Mar 2026
PA3: Policy-Aware Agent Alignment through Chain-of-Thought
This work proposes a multi-stage alignment method that teaches models to recall and apply relevant business policies during chain-of-thought reasoning at inference time, without including the full business policy in-context.
Shubhashis Roy Dipta, Daniel Bis, Kun Zhou et al.
· arXiv.org · 6 citations