Preprint
Aug 2026
Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance
Empirical evaluations show that two LRP-guided decoding methods that steer generation toward trajectories more relevant to cue tokens promote safer responses, supporting cue-token attribution's role in compliance failures.
Or Biton, Tomer Krichli, Itai Allouche et al.
· 0 citations