Skip to content

Author

Hsin-Tai Wu

We have 2 of 20 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

HindSearch: Trajectory-Level Hindsight Critique for Search-Augmented Reinforcement Learning

Search-augmented LM agents are typically trained with a binary exact-match reward, which throws away most of what a failed trajectory tells us about why it failed. We introduce HindSearch, a hindsight self-distillation procedure for GRPO: after each rollout, a frozen judge writes a short critique of every failed trajectory using the gold answer, and the critique supplies an auxiliary on-policy distillation signal on the student's search actions. On the standard seven-benchmark suite with Qwen2.5-3B-Instruct, HindSearch reaches 39.4% average EM, outperforming prior search-RL baselines. Removing the judge's access to the gold answer erases most of the gain, isolating hindsight as the source of the improvement.

Haowei Liu, Jiamian Wang, Hsin-Tai Wu et al. · 0 citations
Jul 2026

ReliableTableQA:How Much Supervision Does Reliability Annotation Need?

This work introduces ReliableTableQA, a framework for training an LLM to annotate the statistical reliability of tabular QA results, and reframe reliability annotation as a data-efficiency problem and delineate precisely when reinforcement fine-tuning does and does not pay off.

Huei-Chung Hu, Hsin-Tai Wu, Koyo Kobayashi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.