This work identifies a factual-salient layer span within LLMs whose derived signal is selectively elevated for factual tokens and exhibits anomalous spikes at hallucination-prone steps, and proposes DescaPE, a decoding framework that leverages internal model signals to suppress hallucination-prone trajectories at infer...
Hayeong Ryu, Jungmin Yun, B. Lim et al.· 0 citations
Experiments on three LRMs show that SaLT-DPO consistently reduces unsafe rates for both reasoning and answer segments while mitigating degradation in benign compliance and preserving general reasoning performance.
Jungmin Yun, Junehyoung Kwon, Hayeong Ryu et al.· 0 citations
ALTSTEER is an inference-time framework that couples selective intervention with refusal-anchored constructive redirection within a single inference pass, and uses an internal refusal-relevant signal to decide when to steer, and applies staged steering to shift generation from refusal-oriented control toward constructi...
This survey examines the problem as narrative consistency, defined as the task-conditioned preservation of binding propositions in the operative narrative state, and introduces a four-category, fourteen-subtype taxonomy comprising World and Setting, Character-Agentive, Event-Structural, and Narration and Discourse cate...
Keunhyeung Park, Seunguk Yu, Jinhee Jang et al.· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.