Decoding-Level Taboo is introduced, a zero-prompt diagnostic stress test that intervenes directly in logit space at runtime, forcing models out of their nominal paths by dynamically masking primary candidate tokens at word boundaries, forcing machine circumlocution.
Abstract
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generation corridor. In real-world deployments, however, complex system prompts, safety guardrails, and structural constraints continuously force models off this nominal path, driving a divergence between benchmark scores and deployment performance. To address this issue, we introduce Decoding-Level Taboo, a zero-prompt diagnostic stress test that intervenes directly in logit space at runtime, forcing models out of their nominal paths. By dynamically masking primary candidate tokens at word boundaries, Taboo forces machine circumlocution. Evaluating Taboo across several open-weight model families reveals that off-path robustness is heavily influenced by both parameter scale and post-training instruction alignment, with robustness generally improving with model size and alignment. Beyond the results presented in this paper, Taboo provides a novel primitive for generating diverse synthetic datasets, stress-testing runtime safety guardrails, and auditing model reliability prior to real-world deployment.
Hunk-Constrained Direct Preference Optimization is introduced, a training framework that unifies security hardening and functional correction in large language models and demonstrates that HPO achieves substantial security improvements—up to 28 percentage points—while preserving or enhancing functional correctness.
Qian-Shuo Huang, Xin Yin, Xin-Rui Li et al.· ACM Transactions on Software...· 0 citations
SkillSafe-Bench is introduced, a controlled benchmark that scores skill-merged models on static refusal, adaptive jailbreak robustness, and capability retention under a conservative two-judge AND rule, and the static effect of merging is base-conditional.
Large language models (LLMs) have shown promise in automated unit test generation, yet the effectiveness of prompt engineering for small, locally-deployed open-source models remains poorly understood. Following growing interest in local LLM deployment to mitigate data exposure risks, this paper presents a controlled em...
M. Tran, Khang Mai· International Conference on...· 0 citations
Large language models (LLMs) have demonstrated remarkable capabilities across languages, yet their safety
alignment remains predominantly evaluated in monolingual, especially English, settings. Code-switching (also known as codemixing)—the alternation between two or more languages within a single utterance or conversat...
Pulagam Naveen Kumar, Sowjanya Bojja, K. Anoosha et al.· International Journal for Re...· 0 citations
This work presents the first systematic characterization of SDC vulnerability across major computation interfaces in both the forward and backward passes of Transformer training, and proposes TrainSDC, a characterization-guided protection framework consisting of Q/K-path recomputation, residual-gain monitoring, and exp...
Zhijie Xia, Haotian Xu, Si-Yu Yun et al.· 0 citations
Code translation, as a challenging and fundamental task, is increasingly relying on large language models (LLMs). However, LLMs often give seemingly plausible but fallacious translations, misleading and even deceptive to debugging developers. We propose tHinter, an automated approach that frames translation error local...
Shengnan Wu, Xin-Yu Sun, Xin Wang et al.· ACM Transactions on Software...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.