Skip to content

Author

Weibo Gao

University of Science and Technology of China

We have 6 of 52 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

Offline Guidance, Online Reasoning: Reusing LLM Feedback for Small Language Models

Reusable Latent Correction (RLC) is proposed, which converts one-off natural-language guidance from a black-box LLM into persistent corrective experiences in the hidden space of an SLM, enabling the SLM to reuse LLM-derived corrections during inference without any online LLM calls.

Bo-Han Zhang, Li-Nan Yue, Weibo Gao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Learner Behavior to Reusable Skills for Effective and Efficient Learner Simulation

Learner simulation aims to reproduce how a particular learner behaves on new tasks. Although Large Language Models (LLMs) can generate increasingly fine-grained learning behaviors, existing approaches often need to repeatedly process a growing interaction history to reconstruct the learner. This introduces additional c...

Zi-Jian Chen, Zheng Zhang, Miao Jia et al. · 0 citations
Book Open access Aug 2026

Find Tailored Step Example for Next Step: a Targeted Step-wise Retrieval Framework for Guiding LLM Reasoning

This work proposes Step-wise Training for In-context Reasoning (STIR), a model to dynamically decide when to retrieve a single logically consistent next step, just using the current problem and its intermediate state as the query.

Cheng Yang, Zhenya Huang, Liyang He et al. · 0 citations
#artificial intelligence Preprint Sep 2026

UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation Paradigms

Reinforcement learning (RL) excels on tasks with verifiable rewards, but in open-ended tasks, the reliability of reward models remains a key challenge. Existing solutions either depend on costly proprietary LLM-as-a-Judge systems or opaque scalar reward models that lack interpretability. Recent works on generative rewa...

Peng Lai, Yi-Chao Du, Junchao Wu et al. · 1 citation
Preprint Jul 2026

AI-generated Images Challenge Visual Trust in High-risk Scenarios

SafeIMG is introduced, a safety-oriented benchmark spanning 12 public- and individual-safety scenarios generated using GPT Image 2.0 that provides human annotations that localise suspicious regions and explain local artefacts and higher-level commonsense or physical inconsistencies.

Yizhi Wang, Yichen Xiao, Linan Yue et al. · 0 citations
Book Open access Aug 2026

Find Tailored Step Example for Next Step: a Targeted Step-wise Retrieval Framework for Guiding LLM Reasoning

This work proposes Step-wise Training for In-context Reasoning (STIR), a model to dynamically decide when to retrieve a single logically consistent next step, just using the current problem and its intermediate state as the query.

Cheng Yang, Zhenya Huang, Liyang He et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.