Skip to content

Author

Jiaxin Mao

We have 5 of 17 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Oct 2026

Learning to Retrieve via Reinforcement Learning in Embedding Space

Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables...

Qi Liu, Feng-Ming Liang, Yi-Qun Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation

Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuff...

Zi-Xuan Yang, Yi-Qun Chen, Qi Liu et al. · 0 citations
Preprint Aug 2026

Diagnosing Search Behavior and Failure Modes in Long-Horizon Search Agents

Deep search agents answer difficult information-seeking questions by iteratively issuing search queries to gather supporting evidence, but it remains unclear whether and how greater search effort leads to better answers. We study these questions through a trajectory-level diagnosis of long-horizon search agents. Using...

Qi Liu, Jiaxin Mao, Fengbin Zhu et al. · 0 citations
Preprint Aug 2026

Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents

Fetch-then-Explore is proposed, which separates page selection from evidence extraction and keeps what it selects: pages are recorded in a per-question workspace on the filesystem rather than the context window or a transient session, and evidence is pulled from them on demand later.

Qi Liu, Yi-Qun Chen, Zidan Chen et al. · 0 citations
Jul 2026

SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task

SearchArt is introduced, a scalable framework for training long-horizon search agents through verification-driven task synthesis and a multi-stage post-training pipeline, which exhibits adaptive search planning, iterative evidence aggregation, and complex reasoning over extended interaction horizons.

Lang Mei, Xiao-Han Yu, Chong Chen et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.