Skip to content

Author

Jian Luan

We have 9 of 71 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Oct 2026

AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning

Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of individual decisions. Step-level credit assignment provides finer-grained supervision, but its estimates can be unreliable because observed returns also depend on s...

X. Wang, Wen-Hao Wu, Meng-Hao Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

HarnessPAI: An Evolving Harness for Physical AI

The results suggest that the frontier of Physical AI depends not only on stronger action models, but also on executable harnesses that integrate perception, task understanding and reasoning, and action execution into a unified, verifiable, and feedback-driven system.

X. Wang, Wen-Hao Wu, Meng-Hao Zhang et al. · 1 citation
#natural language process... Preprint Sep 2026

Where to Look and What to Use: Retrieve-Localize-Generate for Long-Term Conversational Memory Question Answering

This work proposes MemLoc, a unified Retrieve-Localize-Generate framework for long-term conversational memory QA, and introduces a reasoning-based evidence locator trained with Self-reflective Hint Policy Optimization, which performs progressive refinement by extracting query-relevant fragments within memory units to s...

Yi-Fan Wang, Xin-Kui Lin, Yong-Xiu Xu et al. · 0 citations

GUI-PRA: Process Reward Agent for GUI Tasks

gui-PRA is introduced, a Process Reward Agent that transforms GUI process evaluation from passive scoring into active investigation, and demonstrates strong competitiveness against fully trained critic models.

Tao Xiong, Xavier Hu, Yurun Chen et al. · 5 citations
Preprint Aug 2026

TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents

Results show that TRACE converts high model potential into stable, consistent performance gain, and bridge the gap between potential and reliable performance to just 4.0 points.

Wen-Hao Wu, Meng-Hao Zhang, X. Wang et al. · 2 citations
Preprint Aug 2026

Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

This work studies reference-free post-training for multilingual machine translation with open large language models and finds that on-policy distillation reaches, but does not surpass, the quality frontier achieved by RL with checkpoint interpolation.

Chris Han, Pengzhi Gao, Pei Fu et al. · 0 citations
#machine learning Preprint Jul 2026

SEE: Structure-aware Exploring&Exploiting for Long-horizon GUI Agent Trajectory Synthesis

See, a two-stage data synthesis framework consisting of an efficient exploration stage that builds an explicit UI transition graph over screens and elements, and a graph-based synthesis stage that composes diverse multi-step trajectories via planning and controlled sampling, yields reproducible and explainable data gen...

Zhuohang Fan, Beichen Zhang, Yuanfa Li et al. · 0 citations
#artificial intelligence Preprint Aug 2026

G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution

G-ReAct is a reasoning framework for deep search that organizes reasoning as state evolution over a fixed-topology query graph, transforming exploratory search driven by textual history into graph-guided reasoning under explicit constraints.

Shaoxiong Yang, Mengyuan Zhang, Shao-Jun Lin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.