Skip to content

Category

machine learning

12,457 papers

#machine learning Preprint Oct 2026

Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects

Large language models (LLMs) are increasingly used as judges for automated AI evaluation. A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear. We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear...

Tian-Xi Li, Jie Ding · 0 citations
#machine learning Preprint Open access Oct 2026

High-dimensional online calibration from harmonic weights

We study the online calibration of multidimensional forecasts over an arbitrary convex set $Y\subseteq\mathbb{R}^d$ relative to an arbitrary error norm $\|\cdot\|_{L}$. For forecasting $d$ binary outcomes simultaneously ($Y=[0,1]^d$), we give the first algorithm that achieves $\varepsilon$-calibration in a number of ro...

Maxwell Fishelson, Mehryar Mohri · 0 citations
#machine learning Preprint Open access Oct 2026

Nash Social Welfare for Multi Armed Bandits: Trajectory-wise Expected and High Probability Regret

We study fair multi-armed bandits under the Nash Social Welfare (NSW) objective, which measures performance via the geometric mean of accumulated rewards. Existing work defines Nash regret as $\mathrm{NR}_T = \mu^\star - (\prod_{t=1}^T \mathbb{E}\mu_{I_t})^{1/T}$, where $\mu_{I_t}$ is the mean reward of the recommended...

Avishek Ghosh · 0 citations
#artificial intelligence Preprint Oct 2026

Learning to Retrieve via Reinforcement Learning in Embedding Space

Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables...

Qi Liu, Feng-Ming Liang, Yi-Qun Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SanSi: A Looped Typed Decision Model for System 1.5 Thinking

Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass. A single pass is fast, intuitive System 1 thinking. We study what lies between one pass and generated reasoning: looping, in which the same layers ar...

Shuyu Gan, Young-Jun Lee, Dongyeop Kang · 0 citations
#machine learning Preprint Open access Oct 2026

The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models

Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-si...

Yibo Zhang, Tianrong Guan, Liang Lin et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Exact Calibration and Sharp Risk Geometry for Volume-Sampled Ridge Regression

We study ridge regression from exactly $s$ distinct rows of a fixed design. Responses are fixed, and only the subset is random. The determinant law and selected ridge fit share one positive definite penalty. Established mean identities and exponential-family duality give the unique penalty that matches a prescribed ful...

Kihun Rhee · 0 citations
#machine learning Preprint Open access Oct 2026

RefRoute: Decoupling Conditioning Cost from References via Compact Residual Conditioning and Spatial Routing

Multi-reference image generation requires preserving the appearance of multiple subjects while composing them into a coherent scene. However, existing diffusion transformers commonly encode references as dense visual token grids and jointly process them with global attention, making conditioning increasingly expensive...

Wanning He, Yuyao Zhang, Yu-Wing Tai · 0 citations
#machine learning Preprint Open access Oct 2026

Stability of Measure-to-Measure Transformers on Sub-Gaussian Data

Transformers have exhibited impressive empirical success across various domains, but their theoretical foundations remain less developed. This work constitutes a mathematical study of the measure-to-measure operators defined by transformers. We show that transformers map sub-Gaussian inputs to sub-Gaussian outputs; thi...

Frank Cole, Nicholas H. Nelsen, Takashi Furuya · 0 citations
#machine learning Preprint Oct 2026

Mathematical Invariant-Enabled Topological Neural Networks for Molecular and Materials Property Prediction

Existing molecular and materials learning approaches often rely on a limited set of structural representations, which may capture only selected aspects of complex three-dimensional structure. Here, we introduce mathematical invariant-enabled topological neural networks (MITNNs), a framework that represents complex stru...

Yi-Ming Ren, Xiang Liu, Mustafa Hajij et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models

Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication. Such a stringent information structure renders the conventional reward...

Fernando Martinez, Tao Li, Yingdong Lu et al. · 0 citations
#artificial intelligence Preprint Oct 2026

CACHEFORGE: LLM-Guided End-to-End Generative Cache Replacement Policy for Performance and Hardware Efficiency

Modern cache replacement designs saturate because they operate within fixed representational structures, hand-crafted and heuristic based feature-engineered predictors, or offline imitation models that cannot generate new decision logic on their own. At the same time, replacement is shaped by the causal interaction of...

Kaushal Mhapsekar, Bita Aslrousta, Brijesh Kumar Bhayana et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.