Skip to content

Author

Yushi Sun

We have 6 of 26 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents

Audio language models can recognize spoken commands and invoke tools, but an agent must first decide whether the acoustic and conversational context warrants action. We introduce VGBench, a 1,018-item diagnostic benchmark for action-level addressedness across side-talk, self-talk, and speaker-switch scenarios. Each ite...

Yan-Jie Zhang, Nan-Chen Hu, Yu-Shi Sun · 0 citations
#artificial intelligence Preprint Sep 2026

When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents

Intent drift is established as a measurable multi-turn failure mode and explicit state maintenance as a partial mitigation and IntentFlux is introduced, an executable benchmark that converts verifiable tasks into dialogues with controlled intent changes while preserving their original graders.

Yan-Jie Zhang, Bo-Wen Cao, Zi-Xin Chen et al. · 0 citations
Preprint Aug 2026

When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents

Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the environment changes. We ask what happens when an agent must reconcile a confident memory claim with a contradicting observation, and whether current models can catch the conflict before it becomes a safety-rel...

Yu-Shi Sun, Yan-Jie Zhang · 1 citation · ⚡1
Preprint Aug 2026

MemoryCPT: An End-to-End Agent Memory Framework for Cost-Performance Trade-off

Experiments on LoCoMo and LongMemEval show that MemoryCPT improves the cost-performance trade-off over the evaluated baselines, while ablation and sensitivity analyses characterize the contributions of its components and the effects of key design choices.

Songxin Lei, Kun Ouyang, Wei-Lin Ruan et al. · 1 citation · ⚡1
Jul 2026

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

SDABench is introduced, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics).

Chuhan Shi, Xiaoquan Ren, Sicheng Song et al. · 1 citation
Preprint Aug 2026

MolecularCanvas: LLM-assisted Small-Molecule Drug Discovery via Structure-Guided Constraints

MolecularCanvas is an interactive system that enables users to iteratively construct an optimization context by integrating high-level goals, structure-level annotations, property constraints, and reference-based preferences that guides the generation of candidate molecules across diverse molecular structures.

Haoyu Dong, Rui Sheng, Shu-Hao Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.