Skip to content

Author

Kui Ren

We have 4 of 67 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Oct 2026

ASCENT: First-Order Optimal Fine-Tuning with Recalibration for Safety--Utility Co-Enhancement

Supervised fine-tuning can substantially improve the downstream utility of large language models (LLMs) but may compromise their safety. Existing safety-preserving methods constrain downstream updates using safety-related parameters or subspaces, but mainly focus on safety preservation rather than joint safety and util...

Wei-Wei Qi, Chong-Yu Wang, Tian-Hang Zheng et al. · 0 citations
Preprint Aug 2026

ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents

Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavioral safety and introduce ActBench, a self-evolving benchmark that evaluates such behavior risk from execution trajectories rather than final responses. Each case pairs a...

H. Yao, Yimin Liu, Meihui Chen et al. · 0 citations
Jul 2026

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection

DARWIN, an evolutionary attack-defense framework that models jailbreaking as a continual process and updates guardrails through an attack-defense loop is proposed, an evolutionary attack-defense framework that models jailbreaking as a continual process and updates guardrails through an attack-defense loop.

Weiwei Qi, Ze-Feng Wu, Zhiling Guo et al. · 4 citations
Preprint Jul 2026

DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment

DataShield is a data assessment framework that identifies risky fine-tuning samples and response segments through consensus subspace alignment over joint safety-critical semantic spaces derived from multiple safety-aligned LLMs, allowing both sample-level filtering and fine-grained segment-level masking.

Ze-Feng Wu, Weiwei Qi, Jielong Chen et al. · 4 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.