Skip to content

Author

Zhongxiang Dai

We have 7 of 23 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Oct 2026

GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution

The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, automatically optimizing this harness poses three coupled challenges: reconciling model inten...

Ge-Yi Yang, Zi-Kun Qu, Xiang Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning

Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Un...

Zi-Hao Chen, Fan-Xiang Xiong, Hong-Ran Ren et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Preference to Reciprocity: Decentralized Matching with Empirically Grounded LLM-agent Based Modeling

Bipartite matching is a fundamental problem in game theory and market design. Classical approaches such as Gale--Shapley assume complete preferences and centralized computation, whereas many real-world matching processes are decentralized, asynchronous, and shaped by sequential interaction under limited information. We...

Wang-Xuan Fan, Xiao-Yu Nie, Zhou-Tian Shi et al. · 0 citations
Book Open access Aug 2026

CES: Combinatorial Experts Selection via Contextual Linear Bandits

With the rapid advancement of large language models (LLMs), multi-agent systems have emerged as a promising alternative to scaling up a single model. Existing approaches ensemble multiple LLMs to improve response quality, but they often rely on static prior knowledge of model capabilities and prompts, and require exten...

Jinkun Xu, Minghan Wang, Zhiyong Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

COBRA-Skills is introduced, an efficient framework that formulates skill optimization as budgeted sequential optimization over a dynamically evolving candidate space and remains robust to changes in the agent harness and performs effectively when the target model itself is used for skill generation and refinement.

Ping-Chen Lu, Xiang-Yi Wang, Xiang Li et al. · 0 citations
Preprint Aug 2026

SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation

Sparse Probing and Outcome-calibrated Targets OPD is introduced, which addresses two coupled decisions, where to probe and what to distill, through an acquisition--exploration--exploitation procedure and demonstrates the effectiveness of SPOT in improving reasoning performance while balancing solution quality and cover...

Zi-Kun Qu, Min Zhang, Ming-Ze Kong et al. · 8 citations

Meta-Prompt Optimization for LLM-Based Sequential Decision Making

The EXPonential-weight algorithm for prompt Optimization} (EXPO) is proposed to automatically optimize the task description and meta-instruction in the meta-prompt for LLM-based agents and is extended to additionally optimize the exemplars (i.e., history of interactions) in the meta-prompt to further enhance the perfor...

Ming-Ze Kong, Zhiyong Wang, Yao Shu et al. · 7 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.