Skip to content

Author

Yuling Shi

We have 12 of 59 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Oct 2026

Datura: Progressive Red Teaming Testing for Tool Invocation Chain in LLM Agents

LLM agents that invoke external tools face critical safety vulnerabilities when malicious manipulations exploit their implicit trust in tool outputs and metadata. However, identifying these vulnerabilities through testing is challenging due to the need to bypass safety guardrails with semantically legitimate inputs, th...

Yu-Chen Shao, Zi-Qun Bao, Yu-Heng Huang et al. · 0 citations
Book Open access Jul 2026

DLLM-Searcher: Adapting Diffusion Language Model for Efficient Search Agents

Recently, Diffusion Large Language Models (dLLMs) have demonstrated unique efficiency advantages, enabled by their inherently parallel decoding mechanism and flexible generation paradigm. Meanwhile, despite the rapid advancement of Search Agents, their practical deployment is constrained by a fundamental limitation, te...

Jiahao Zhao, Shaoxuan Xu, ZhongXiang Sun et al. · 0 citations
#natural language process... Preprint Sep 2026

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

This work introduces early outcome prediction, a complementary axis of efficiency that instead cuts cost within each task within each task, and instantiates EarlyEval, a lightweight framework that trains a pair of LightGBM success and failure classifiers over behavioral, textual, and reference-solution features and hal...

Yu-Ling Shi, Zhensu Sun, Jun-Sen Dong et al. · 2 citations · ⚡1
#natural language process... Preprint Sep 2026

Membership Inference in Fine-tuned Diffusion Language Models via Token-level Memorization Asymmetry

This work identifies a phenomenon termed token-level memorization asymmetry through theoretical analysis of diffusion training dynamics and proposes Q-Skew, a quantile-weighted skewness-based indicator for membership inference on finetuned DLMs.

Shengfang Zhai, Leo Marchyok, Yu-Ling Shi et al. · 0 citations
Preprint Aug 2026

DeepRepoQA: Code Repository Question Answering with Deep Agent Exploration

DeepRepoQA is proposed, a novel question answering (QA) framework for repository-level code understanding that builds on an agentic framework where LLM agents find answers through a systematic tree search over the repository structure.

Wei Peng, Yu-Ling Shi, Yingwei Ma et al. · 0 citations
Jul 2026

Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution

LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone to factual errors caused by insufficient repository understanding. Recent methods attempt to mitigate this limitation through pre-repair repository exploration; however, their fix-driven strategies expl...

Hao-Tian Lin, Si-Lin Chen, Xiao-Dong Gu et al. · 2 citations
Preprint Aug 2026

ParaTempo: Efficient Parallel Reasoning via Temporal Confidence

ParaTempo is a training-free asynchronous parallel reasoning framework driven by temporal confidence, a branch-local measure of answer-space convergence that reduces average latency, and exhibits stronger temporal stability and predictive power for future branch convergence than token-level and instantaneous signals.

Xuteng Zhang, Wenhao Zeng, Xiao-Dong Gu et al. · 1 citation
Preprint Jul 2026

SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

SWE-Pruner Pro is proposed, which prunes tool outputs directly inside the agent, with a small head turns the agent's own internal representations into a keep-or-prune label for each line, with a length-aware embedding keyed to each tool output's line count.

Yuhang Wang, Yuling Shi, Shaoqiu Zhang et al. · 2 citations
Review Aug 2026

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

SWE-Bench ProMax is introduced, an expert-curated, multilingual code refactoring benchmark of 170 instances drawn from real commits across seven programming languages, which presents a meaningful and unsaturated challenge for current AI coding agents.

Yu-Ling Shi, Jing-Heng Xu, Kelin Fu et al. · 6 citations
Preprint Aug 2026

SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

Self-Reflective Policy Optimization (SRPO) enables LLMs to analyze their own completed trajectories, synthesize errors into concise"reflection patches," and use reflection-conditioned teacher scores on student on-policy rollouts as dense token-level training signals.

Jialong Liu, Yu-Ling Shi, Ning Yang et al. · 1 citation
#software testing Preprint Aug 2026

Repo0: Design-Driven Zero-to-All Code Generation

Repo0 is presented, a continuous structural evolution framework for zero-to-all code generation that maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation.

Si-Lin Chen, Haoyi Teng, Xiao-Dong Gu et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.