Skip to content

Fine-tuning as Jailbreaking: A data-centric red teaming framework via logic injection

Aug 2026 · International Conference on Automated Software Engineering · Vol 33 · 0 citations · 31 references

TL;DR

A red-teaming testing method for fine-tuning-stage defenses that audit the harmful content of poisoned fine-tuning data, exposing a vulnerability of the RFT data supply chain to logic injection and point to the need for fine-tuning-stage defenses that audit the harmful content of poisoned fine-tuning data.

View source

Similar papers

Conference Open access 2026

LogSanitizer: Defending LLM-Integrated SOCs against Backdoor Triggers Delivered through Firewall Logs

LogSanitizer is proposed, a family of input sanitization defenses operating at two levels: a pre-prompt log-transformation pipeline that disrupts trigger patterns in the structured log representation, and a post-tokenizer perturbation strategy that corrupts trigger-bearing token configurations before they reach the model.

Leszek Wronski, Bogdan Ksiezopolski · 0 citations
Jul 2026

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

This work proposes Routing-based On-Policy Distillation (ROPD), a novel realignment framework that models the divergence between aligned and compromised output probability distributions rather than fitting specific prompt templates, establishing a new standard for robust LLM realignment.

Yongjian Guo, Wanlun Ma, Lingyu Shen et al. · 0 citations
#small language model Preprint Aug 2026

RTLGuard: A Lightweight Teacher-Student Defense for Poisoned RTL Code Generation Models

RTLGuard leverages a teacher-student framework designed to sanitize compromised RTL generation models by fine-tuning a small-scale,"clean"teacher model on a limited set of trusted RTL data, and incorporating feature alignment and knowledge distillation to suppress malicious behaviors.

Mahshid Rezakhani, K. Azar, H. Kamali · 0 citations
Book Open access Aug 2026

The 2nd SeT-LLM Workshop on Secure and Trustworthy Large Language Models

Large language models (LLMs) are increasingly embedded as core components of data-centric systems, supporting analytical decision making, and automated reasoning over large-scale, heterogeneous datasets. Yet their deployment in open-world environments raises fundamental challenges to security and trustworthiness: LLMs can leak sensitive data, fall prey to prompt injection and jailbreaks, generate misinformation, and behave unpredictably under adversarial inputs, failures that propagate through data pipelines and affect downstream decisions. The rise of LLM-based agents further amplifies these risks through unsafe tool use and autonomous decision-making. The 2nd SeT-LLM Workshop on Secure and Trustworthy Large Language Models brings together researchers and practitioners from data mining, machine learning, security, and responsible AI to address these issues from a data-centric, system-level perspective, spanning robust defenses, trustworthy evaluation, privacy and copyright protection, robustness, alignment and safety, agent security, and high-stakes applications. Through invited talks, contributed papers, a poster session, and a panel discussion, the workshop prioritizes early-stage ideas, system experiences, and open problems across the lifecycle of LLM-based systems.

Lu Lin, Jinghui Chen, Ting Wang et al. · 0 citations
Jul 2026

Just Testing, Move Along: Evasion of LLM-based System Log Interpretation by Prompt Injection

This paper presents a framework for evaluating prompt injection attacks against LLM-based log interpretation using log traces generated during real cyber attacks, and creates adversarial examples through generic injection generation, refinement, and attack-specific optimization.

Max Landauer, F. Skopik, Markus Wurzenberger et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.