Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model wei...
Yen-Jen Wang, Hao-Zhe Jiang, Shu-Ying Deng et al.· 0 citations
When inference demand exceeds available compute capacity, model providers must decide which requests should be served first. Users have different tolerances for delay from an LLM API, but current priority pricing schemes compress these differences into coarse fixed-price service tiers. We design an inference auction th...
Keegan Harris, Siddharth Prasad, Asher Trockman et al.· 0 citations
The results strictly weaken the assumptions required by prior work in the multi-agent information aggregation literature, filling a gap that had remained elusive even for games with constant $CC_\alpha(G)$.
Mark Bedaywi, Scott Emmons, Nika Haghtalab et al.· 0 citations
Large Language Models (LLMs) have achieved remarkable reasoning capabilities by utilizing chain-of-thought (CoT) as a scratchpad for intermediate stages of thinking. However, CoT techniques require explicit supervision on thinking tokens, which requires rich, task-specific data. In this work, we propose Abstract Token...
Khashayar Gatmiry, Avrajit Ghosh, Parsa Mirtaheri et al.· 0 citations
While Reinforcement Learning from Human Feedback (RLHF) is the standard paradigm for aligning large language models with human preferences, its effectiveness in pluralistic settings has been called into question. Notably, recent work by G\"olz et al. (2025) demonstrated that the \textit{distortion} -- defined as the mu...
The notion of assistance regret is introduced: the gap between the cumulative utility of interactions and that of the optimal joint policies in hindsight, which map latent states to action pairs, is introduced.
Nivasini Ananthakrishnan, Mark Bedaywi, Michael I. Jordan et al.· arXiv.org· 0 citations
This work proposes a game-theoretic framework that gives this reward-retention trade-off an explicit statistical interpretation, and provides a principled method for learning this equilibrium coefficient via reduction to the KL-regularized RL objective, thus allowing for flexible integration into standard fine-tuning p...
Keegan Harris, Brian Lee, Ian Waudby-Smith et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.