Skip to content

Author

Xiaoyong Wei

We have 3 of 15 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Meta-Moderator: Empowering Multi-Agent Debate with Meta-Cognition

Multi-agent debate can improve large language model reasoning by eliciting diverse hypotheses and critiques, yet its performance is often constrained by weak moderation. Common pipelines rely on fixed budgets, agreement-based stopping, or untrained judges, leading to redundant deliberation and unreliable evidence aggregation. We cast moderation as a meta-cognitive process, monitoring debate utility, controlling deliberation, and adjudicating a final answer, and introduce Meta-Moderator, a learnable framework that dynamically regulates debate and decides when to finalize an answer. Meta-Moderator is trained independently of the debaters via outcome-driven policy optimization, making debate regulation an explicit capability rather than an incidental effect of prompting. Across five benchmarks, Meta-Moderator outperforms widely used decision layers and transfers across tasks and system configurations. Further analyses show that it allocates debate more selectively and reduces mis-aggregation after informative hypotheses appear.

Wentao Hu, Zhuoyue Wan, Jinhao Shen et al. · 0 citations
Preprint Jul 2026

Explaining and Tuning Transformer-based LLMs in Arithmetic Tasks with Human Strategies

The findings reveal that transformer-based large language models exhibit learning patterns similar to those of human learners, with a faster learning speed for simpler subtasks compared to more complex ones, which suggests that transformer-based LLMs may share cognitive processes with human learners in arithmetic.

Luyu Qiu, Jianing Li, Hwanhee Kim et al. · 0 citations

Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad

The results show that chemical CoT is neither a faithful explanation nor merely a post-hoc rationalization, but a hallucination-prone molecular scratchpad, which cautions against treating CoT as direct evidence of faithful reasoning and motivates process-level supervision beyond answer-only evaluation.

Jiatong Li, Yuxuan Ren, Weida Wang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.