Skip to content

Category

machine learning

12,457 papers

#machine learning Preprint Open access Oct 2026

Reflected Anchored Langevin Algorithms

First order Langevin algorithms for constrained sampling in machine learning, such as projected Langevin Monte Carlo which are based on discretizations of reflected Langevin dynamics, require differentiable log densities that limits their applicability. This paper introduces reflected anchored Langevin dynamics (RALD),...

Changwei Tu, Xiaoyu Wang, Yingli Wang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

It Is Not Seeing the Hazard: A Frozen Vision-Language Safety Score Measures Its Caption Bank

Frozen vision-language models increasingly provide safety signals for reinforcement learning. Their use assumes that similarity to language describing danger indicates the hazard itself. Yet policy return and collision rate cannot reveal whether a score detects hazards or responds to correlated features of the scene. V...

Samuel Tetteh, Cody Fleming · 0 citations
#machine learning Preprint Open access Oct 2026

TERRA: Learning Transportable Latent Actions through Temporal Effect Representation and Relational Alignment

Latent actions supervise robot policies with action-like codes inferred from visual transitions, and their usefulness hinges on two questions: what a code keeps from a transition, and whether it still means the same thing when reused in a different initial state. The first is a tension in time: an endpoint difference d...

Tianxingjian Ding, Mubarak Shah, Yu Tian · 0 citations
#machine learning Preprint Open access Oct 2026

Adjoint-Based Calibration and Optimal Control of Stochastic Multiscale Bioprocess Digital Twins

We develop a bias-aware digital-twin calibration and control framework for multiscale bioprocess models within a biological systems-of-systems (Bio-SoS) paradigm. The digital twin is represented by a stochastic differential equation (SDE) model and calibrated from sparse, discrete observations using quasi-likelihood es...

Keilung Choy, Wei Xie · 0 citations
#machine learning Preprint Open access Oct 2026

An Invariant Tangent-Angle Descriptor and a Band U-Net for 2D Fragment Adjacency Prediction

This paper addresses the prediction of adjacency between pairs of 2D fragments based on their contours. We improved the two-stage architecture proposed in Beaulac's thesis, in which a rotation-equivariant Siamese convolutional neural network scores pairs of local image windows along the two contours of two fragments. T...

Guillaume Brouillette (Universit\'e du Qu\'ebec \`a Trois-Rivi\`eres, Trois-Rivi\`eres, Canada) et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Democratizing MoE inference on commodity GPUs with CoMoE

Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication. Consequently, state-of-the-art inference systems require high-bandwidth, P2P interconnects (e.g., NVLink) in datacenter GPUs to handle massive token routing, making deployment prohibitively e...

Ruwen Fan (Jimmy), Yuezhi Zu (Jimmy), Junru Li (Jimmy) et al. · 0 citations
#machine learning Preprint Open access Oct 2026

The Persona Hierarchy Model: Understanding Contextual Generalization in Fine-Tuning LLMs

Language models are routinely fine-tuned under a fixed context, such as a generic system prompt, persona or domain-specific instruction, yet the learned behavior sometimes stays confined to that context and sometimes broadly generalizes to unseen contexts. We propose the Persona Hierarchy Model to explain this: a share...

Jiachen Zhao, Zhengxuan Wu, David Bau et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Immiscible Diffusion Policy: Preserving Multimodal Robot Actions through Label-Free Noise Assignment

When diffusion policies were first introduced, they were expected to recover multi-modal action distributions. However, we find this expectation does not always hold, as diffusion policies often collapse to a single modality even when we guarantee the balance of dataset modalities and exact within-batch symmetry. Our a...

Xiao Zhang, Yuxin Chen, Zhixuan Liang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Closing the Loop on Contrail Avoidance with Satellite Verification

Contrails are the thin ice clouds that aircraft leave behind. They cause a large share of aviation's warming, and rerouting the few flights that produce them could avoid much of it. However, an avoided contrail only counts if a satellite can confirm that it never formed, and this check is hard: contrails are one to two...

Spandan Ghose Chowdhury · 0 citations
#machine learning Preprint Open access Oct 2026

Learning Unknown Constraints without Unsafe Data via Optimality and Counterfactual Regularization

Learning from demonstrations (LfD) provides a framework for inferring unknown constraints from locally optimal, constraint-satisfying expert behavior. Existing approaches largely fall into two paradigms, constrained inverse optimal control (CIOC) and inverse constrained reinforcement learning (ICRL). CIOC exploits opti...

Zhouyu Zhang, Chih-Yuan Chiu, Glen Chou · 0 citations
#machine learning Preprint Open access Oct 2026

Ask the Expert: LLM-Guided Reinforcement Learning for Autonomous Cyber Defense

Policy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations with actions from large action spaces. While large language models (LLMs) may reason semanti...

Fernando Martinez, Abhishek Satyam, Tao Li et al. · 0 citations
#machine learning Preprint Open access Oct 2026

VM-ARRAYDPS: Virtual Microphone Augmented Diffusion Posterior Sampling for Unsupervised Blind Speech Separation

Blind Source Separation(BSS) is a fundamental problem in signal processing, aiming to separate multiple source signals from their mixtures without prior knowledge of the sources or the mixing process. Traditional approaches, such as Independent Vector Analysis (IVA) exploits statistical independence of sources. Recentl...

Jingqi Sun, Haozhan Tang, Shulin He et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.