Skip to content

Author

Su Wang

We have 6 of 18 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Noise Floor Audit for Agent Benchmarks

We audit measurement variability for 3 native tool-calling endpoints across 2 providers on the official BFCL multiple and parallel categories, using matched AST grading. At temperature 0, reruns are nearly deterministic across Groq endpoints and a thinking-enabled Gemini setting: ever-flip fractions are 0.7%, 2.0%, and...

Yi-Hang Chen, Pinyan Qian, Su Wang et al. · 0 citations
Preprint Aug 2026

BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL

BAP-SQL is presented, which treats observation formation as a budget-control stage: it estimates query risk, rewrites SQL when useful, and delegates hard limits to an independent runtime shield and improves tight-budget success.

Chong Peng, Pinyan Qian, Su Wang et al. · 2 citations
Jul 2026

Phantom Guardrails: When Self-Improving Agent Harnesses Fix Failures That Never Happened

This work introduces the Counterfactual Fabrication Lab, a deterministic micro-lab where the correct action is known: do nothing, and presents the Counterfactual Fabrication Lab for measuring fabricated failures in self-improving agent harnesses.

Su Wang, Pinyan Qian, Yifan Lin et al. · 6 citations · ⚡2
Preprint Aug 2026

Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification

Family-level and seed-stability analyses show that example-level accuracy overstates counterfactual consistency: complete four-way family success is rare, and an exploratory follow-up that elicits decomposed semantic evidence fails to improve routing for the cleanly evaluated endpoint.

Yi-Hang Chen, Pinyan Qian, Su Wang et al. · 0 citations
Preprint Jul 2026

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions

A minimal benchmark design and candidate reporting metrics for user-conditioned adaptation are proposed and a concrete design requirement for future personal-agent evaluation, with metrics used as reporting tools for that requirement.

Pinyan Qian, Su Wang, Yihang Chen et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.