Skip to content

Author

Xinyuan Song

We have 8 of 21 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

MEMONDEMAND: A Memory Management System for Large-Scale Enterprise Data

On EnterpriseRAG-Bench, MEMONDEMAND outperforms the strongest published LB#1 result at every evaluated scale from 10M tokens through the complete 618M- token collection, and results on FinanceBench, HotpotQA, and FRAMES further show strong performance across financial, multi-hop, and fact-retrieval settings.

Xin-Yuan Song, Bo-Wen Zhu, H. Haque et al. · 0 citations
Preprint Aug 2026

MegaMem: A Retrieval Solution for Ultra-Large Context Windows

These results show that MegaMem supports ultra-large persistent memory while preserving strong answer accuracy under a bounded generation context, and provides a practical path toward accurate retrieval over memories ranging from hundreds of millions to one billion tokens.

Xinyuan Song, Bowen Zhu, H. Haque et al. · 0 citations
Preprint Jun 2026

AlgoBench: Benchmarking Algorithmic Adaptation in Code Generation

AlGOBENCH is introduced, a framework that automatically builds novel algorithmic problems from known competitive-programming problems through structured constraint-shifting transformations and error analysis shows that failures are mainly algorithmic rather than implementation-level, suggesting that ALGOBENCH evaluates adaptation beyond functional correctness.

Xinyuan Song, Z. Cai, Liang Zhao · 0 citations
Preprint Jul 2026

Repair the Amplifier, Not the Symptom: Stable World-Model Correction for Agent Rollouts

WM-SAR, a spectral subgraph repair method that estimates node-edge amplification, greedily grows a connected repair region by marginal residual-spectral relief, and sends only this region to an LLM for root-cause repair, achieves stronger long-horizon stabilization and root-cause recovery under compact token budgets.

Xinyuan Song, Z. Cai · 0 citations
Preprint Jul 2026

Learning to Control LLM Agent Harnesses with Offline Reinforcement Learning

Large language model (LLM) agents are usually improved by changing prompts, models, or hand-written workflows, while the execution harness around the model is treated as fixed infrastructure. We argue that this harness is itself a learnable control layer. We formalize harness operation as a finite-horizon Harness MDP, where a lightweight controller selects structural execution actions while the LLM executor remains frozen. The controller is trained from offline rollouts using advantage-weighted regression with only terminal task-rubric rewards. We also separate final task quality from a post-hoc Harness Maturity Score, which measures whether the harness follows reliable execution patterns rather than only whether the final answer is correct. This separation gives a finite-buffer view of harness learning: final-quality gains require high-return support in the offline buffer, while process behavior can shift whenever it aligns with advantage-weighted actions. Across six controlled domains and two public-benchmark adapters, the learned controller consistently improves verification behavior and selectively improves final task quality, with the largest gains on adapted tau-bench retail, adapted AgentBench DB-Bench, and coding with a calibrated structural verifier. Ablations against behavior cloning and Forced CHECK show that the gains are not explained by imitation or by simply adding checks. These results identify harness control as a learnable layer for frozen LLM agents, while showing that offline support limits when better process control becomes better final answers.

Haiwen Yi, Xinyuan Song · 1 citation
Jun 2026

AlgoSkill: Learning to Design Algorithms by Scheduling Human-Like Skills

Experiments on competitive programming and combinatorial optimization benchmarks show that AlgoSkill improves over direct LLM generation, chain-of-thought prompting, self-refinement, and MCTS without typed skills, which support treating automatic algorithm design as verification-guided skill scheduling rather than one-shot code generation.

Xinyuan Song, Z. Cai, Liang Zhao · 1 citation
Preprint Jul 2026

Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents

A belief-rollout diagnostic is introduced that elicits structured K-step trajectories over progress, risk, recoverability, constraints, failure mode, uncertainty, future success, repair cost, and next action under alternative harnesses and suggests that harness design is an experimental variable in agent evaluation, not an implementation detail.

Haiwen Yi, Xinyuan Song · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.