Jul 2026
Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning
Across controlled poisoning tasks, subliminal learning, and emergent misalignment, consensus decoding suppresses source-specific misbehavior while preserving shared desirable behavior, including cases where union training and weight averaging retain the unwanted behavior.
Adhyyan Narang, Artin Tajdini, C. Zhang et al.
· arXiv.org · 0 citations