Skip to content

Author

Xinran Chen

We have 3 of 80 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

STAMP: Provenance-Guided Credit Assignment for Deep Search Agents

Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware rewards, and evidence coverage. Yet the actions that expose supporting documents receive no targeted credit, a gap we call the reward-credit mismatch. We propose STAMP, in which a reference-based verifier judges whether each cited document supports an entity or relation in a training-time evidence graph, and first-exposure attribution traces each supported citation back to the action that first surfaced it. This step credit is injected through sign-preserving advantage modulation, which redistributes advantage across steps without changing the trajectory-level reward or the relative ranking of trajectories within each group. On BrowseComp, BrowseComp-ZH, and xbench-DS, STAMP improves the GRPO baseline by +2.0/+5.5/+3.0 points under matched SFT initialization, training data, and search tools, and composes with both outcome-only and citation-rubric base rewards. Component ablations confirm that the provenance-based credit signal and the sign-preserving advantage modulation each contribute to the gains.

Ke Xu, Han Xu, Xinran Chen et al. · 1 citation
Jul 2026

SCOPE-RL: Optimizing Reasoning Paths Before and After Success

SCOPE-RL improves average accuracy by up to 11.2 pp and reduces reasoning tokens by up to 27.1% over outcome-only GRPO, indicating that reward-signal densification is complementary to policy-update-level RLVR advances.

Xiaojia Liu, Han Xu, Jianqiang Xia et al. · 0 citations
Preprint Aug 2026

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning

TurnSight is proposed, a turn-level hindsight self-distillation framework that derives supervision directly from execution-conditioned hindsight and selects reliable supervision through cross-horizon directional agreement.

Changle Qu, Sun-Hao Dai, Hengyi Cai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.