InterTab is a structure-aware framework for CoT reasoning over table images that interleaves chain-of-thought with tool calls that crop structure-aligned table regions, and improves the average accuracy of its backbone from 68.28% to 73.17% and achieves the best average performance among all compared methods.
Hanqian Li, Si-Rui Huang, Chen Ling et al.· 0 citations
Agentic retrieval-augmented generation (RAG) requires language models to decide when to continue searching and when to answer. Existing RL-based methods rely on external supervision and overlook the agent's internal belief about whether the current evidence is sufficient. To address this problem, we reformulate the sea...
Qiuyi Qi, Tian Liang, Jiamu Wang et al.· 0 citations
A hierarchical credit assignment framework that retains RL's verifier-bounded ceiling while incorporating dense token-level signals from a privileged self-teacher is introduced, unlocking dense credit assignment without sacrificing the verifier-bounded ceiling.
Zechuan Wang, Siyuan Lu, Hongxuan Zhang et al.· 4 citations
An underlying mechanism explaining the gap in uncertainty quantification is identified: path switching, where agents frequently abandon their current search direction in-trajectory, breaking the link between early signal and final outcome.
Zongrun Li, Cheng-Yue Yu, Lei Zang et al.· 2 citations
A reinforcement learning framework that converts rubric-level functional feedback into localized optimization signals over generated code and aligns evaluator-generated textual attributions with responsible code spans and generated tokens is proposed.
Rui Jin, Jikai Chen, Yihang Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.