Skip to content

Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning

Jul 2026 · arXiv.org · Vol abs/2607.23394 · 0 citations · 45 references
Computer Science

TL;DR

Across controlled poisoning tasks, subliminal learning, and emergent misalignment, consensus decoding suppresses source-specific misbehavior while preserving shared desirable behavior, including cases where union training and weight averaging retain the unwanted behavior.

Abstract

Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly benign data can transmit hidden preferences that generalize broadly. Standard defenses, such as data filtering, mixing in harmless data, and regularization, attenuate these effects but do not eliminate them. We instead pursue robustness through redundancy: collecting multiple datasets from different sources and only learning what is common between them. Thus, if only a subset of sources are malicious, the misbehavior will be blocked. In order to implement this defense strategy, we fine-tune a separate reference model on each source's dataset and aggregate their next-token distributions at decoding time. We introduce two consensus decoders: a token-wise minimum, which caps each token at the lowest probability any source assigns, and a base-relative variant, which reverts to the base probability on any token the sources move in opposing directions. We further relax exact agreement to tolerate partial support across sources and different surface expressions of the same intention. Across controlled poisoning tasks, subliminal learning, and emergent misalignment, consensus decoding suppresses source-specific misbehavior while preserving shared desirable behavior, including cases where union training and weight averaging retain the unwanted behavior.

View source

Similar papers

#machine learning Preprint Sep 2026

Privacy-Preserving Split Learning for Federated LLM Fine-Tuning

This work addresses leakage through a learned obfuscate-and-recover scheme that protects participants' private datasets while still allowing an independently deployable model to be trained on the server side, making split-based federated LLM fine-tuning practically viable.

Heng Jin, Chao-Yu Zhang, He-Xuan Yu et al. · 1 citation
Preprint Aug 2026

Mitigating LLM sycophancy with RL-based fine-tuning: Bayesian Truth Serum approach

Large language models (LLMs) frequently exhibit \emph{sycophancy}: they adapt their answers to a user's stated beliefs or preferences instead of reporting what they hold to be true, which lowers factual accuracy and can amplify misinformation. This paper proposes a methodology for mitigating sycophancy that employs the...

Serhii Mytsyk, Yiming Zhang, Vikram Krishnamurthy · 0 citations
#artificial intelligence Preprint Sep 2026

Chopthin-Consensus Power Sampling: A Diversity-Preserving Approach to LLM Decoding

Inference-time power sampling via Sequential Monte Carlo (SMC) can substantially improve large language model (LLM) reasoning without requiring post-training. However, many existing SMC approaches rely on equal-weight resampling, which can aggressively prune low-weight trajectories, discarding potentially correct reaso...

Minoo Ahmadi, Seyedarmin Azizi, Erfan Baghaei Potraghloo et al. · 0 citations
Conference Open access Aug 2026

Mitigating Backdoors via Decoy Shortcuts and Knowledge Decoupling

This work reveals that backdoor behaviors tend to be absorbed by a simpler parallel branch when jointly trained with the main network, and proposes Trapping and Removing (TR), a simple yet effective training-time defense that introduces a lightweight shortcut branch as a "honeypot" to trap backdoor knowledge.

Zixuan Zhu, Rui Wang, Lihua Jing et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.