Skip to content

Author

Xinfeng Li

We have 7 of 86 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Practical Secrets Extraction against Black-box LLMs

Large language models (LLMs) increasingly power autonomous coding agents such as Codex and Claude Code, yet their training corpora may contain confidential credentials exposed in public repositories or collected from private development artifacts, creating risks of memorization and subsequent leakage. Existing extracti...

Shi-Qian Zhao, Si-Wei Jiang, Xin-Feng Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Concealing LLM-Based Multi-Agent Topology via Phantom Structure Injection

Driven by the rapid advancement of large language models (LLMs), LLM-based multi-agent systems (MAS) have emerged as a powerful paradigm for collaborative reasoning over complex tasks. A key design element of MAS is the communication topology, which governs information flow among agents and often encodes proprietary kn...

Long-Zhu He, Ze-Kun Wen, Xin-Feng Li et al. · 0 citations
#artificial intelligence Preprint Dec 2024

Easier Said Than Done: Unpacking Intent-Behavior Gap in Jailbreaking LLM-Based Robots

This paper introduces POEF (POlicy EFfective Jailbreak), an automated red-teaming framework that takes into account the robot-specific constraints during both the optimization and evaluation processes and proposes two defense strategies that mitigate the behavior jailbreak risks.

Xuancun Lu, Zhen Huang, Xin-Feng Li et al. · 17 citations · ⚡5
Preprint Aug 2026

ECLIPSE: Self-Evolving Stealthy Prompt Injection Attack against Long-Horizon Agentic Systems

ECLIPSE is a self-evolving and stealthy prompt-injection framework for long-horizon agentic systems that achieves up to 96.7% attack success without defense and 69.2% under the common safety filter, exceeding the strongest baseline by 27.5% in the defended setting.

Shi-Qian Zhao, Yang-Fan Zhou, Xin-Feng Li et al. · 0 citations
Preprint Aug 2026

ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization

It is revealed that contexts with stronger semantic-shift capabilities are more likely to guide models toward recovering harmful meanings and achieving successful jailbreaks, and a black-box context-aware semantic-shift jailbreak framework with Iterative Context Optimization is proposed.

Hujian Zhu, Yi-Hao Huang, Felix Juefei-Xu et al. · 0 citations
2026

Patronus: Safeguarding Text-to-Image Models Against Adversarial Fine-Tuning

Text-to-image (T2I) models can be exploited to produce unsafe images. Existing safety measures, e.g., content moderation or model alignment, can be weakened by adversaries who attempt to restore unsafe generation through model fine-tuning. This paper presents Patronus, a defensive framework that improves T2I models’ re...

Xin-Feng Li, Sheng-Yuan Pang, Jialin Wu et al. · 0 citations
Preprint Aug 2026

Beyond Over-Refusal: Defending Indirect Prompt Injection via Latent Instruction Manifolds

AEGIS (Adaptive Ensemble Guard for Injection Shielding) extracts instruction-sensitive projectors to identify malicious instructions and leverages a Unified Multi-Layer Consensus mechanism that aggregates topologically distinct signals across the network depth.

Jia-Hao Chen, Ruiping Yin, Xin-Feng Li et al. · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.