Skip to content

Author

Tian-Hang Zheng

We have 7 of 25 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Oct 2026

ASCENT: First-Order Optimal Fine-Tuning with Recalibration for Safety--Utility Co-Enhancement

Supervised fine-tuning can substantially improve the downstream utility of large language models (LLMs) but may compromise their safety. Existing safety-preserving methods constrain downstream updates using safety-related parameters or subspaces, but mainly focus on safety preservation rather than joint safety and util...

Wei-Wei Qi, Chong-Yu Wang, Tian-Hang Zheng et al. · 0 citations
Jul 2026

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection

DARWIN, an evolutionary attack-defense framework that models jailbreaking as a continual process and updates guardrails through an attack-defense loop is proposed, an evolutionary attack-defense framework that models jailbreaking as a continual process and updates guardrails through an attack-defense loop.

Weiwei Qi, Ze-Feng Wu, Zhiling Guo et al. · 4 citations
Preprint Aug 2026

JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills

This work introduces \textsc{JailbreakSkill}, a skill-centric framework for scaling automated red-teaming through reusable and continuously evolving attack capabilities, which packages existing attack strategies into modular, agent-ready skills that can be directly reused and adaptively selected across tasks and target...

Xiaoyu Wen, Jiajia Li, Zhida He et al. · 2 citations
Preprint Aug 2026

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

This work proposes the first completely probabilistic architecture-agnostic guardrail to leverage the LLM early output distributional signals for estimating and calibrating the safety probability, thereby enabling early stopping of unsafe ongoing outputs.

Xinzhe Huang, Biwu Yao, Kedong Xiu et al. · 0 citations
Preprint Aug 2026

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning

A Unidirectional Safety Gate (USG) is proposed, instantiated as a Null Space Cubic Layer together with an Inverse Adapter inserted after the final Transformer layer, suggesting that release-time representation-space blocking can raise the cost of malicious downstream adaptation without requiring downstream cooperation.

Yuxuan Huang, Xingyu Zeng, Tian-Hang Zheng et al. · 0 citations
Preprint Jul 2026

DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment

DataShield is a data assessment framework that identifies risky fine-tuning samples and response segments through consensus subspace alignment over joint safety-critical semantic spaces derived from multiple safety-aligned LLMs, allowing both sample-level filtering and fine-grained segment-level masking.

Ze-Feng Wu, Weiwei Qi, Jielong Chen et al. · 4 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.