Skip to content

Category

machine learning

12,457 papers

#machine learning Preprint Open access Oct 2026

The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models

Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations. Yet it remains unclear where and how such associations are computed within VLMs. In this work, we show that VLMs rely on two concurrent mechan...

Kelly Cui, Nikhil Prakash, Shoval Messica et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Deep Time-Series Forecasting in 10 Years: A Survey

Autocorrelation is a common property of time-series, where each observation is dependent on its predecessors. In deep time-series forecasting, it raises two central challenges: (1) designing backbone architectures to model autocorrelation in history sequences, and (2) devising loss functions to model autocorrelation in...

Hao Wang, Licheng Pan, Qingsong Wen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

World Properties without World Models: Distributional Associations and the Interpretation of Decoding Results from Language Models

A growing literature shows that variables can be linearly decoded from the activations of large language models (LLMs). These range from properties of the world, such as the locations of cities and the lifetimes of historical figures, to emotions and pain. Such findings are often taken as evidence that language models...

Elan Barenholtz · 0 citations
#machine learning Preprint Open access Oct 2026

[b] = [d] - [t] + [p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic

Self-supervised speech models (S3Ms) are known to encode rich phonetic information, yet how this information is structured remains underexplored. We conduct a comprehensive study across 96 languages to analyze the underlying structure of S3M representations, with particular attention to phonological vectors. We first s...

Kwanghee Choi, Eunjung Yeo, Cheol Jun Cho et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Uncovering Cross-Objective Interference in Multi-Objective Alignment

We study a persistent failure mode in multi-objective alignment for large language models (LLMs), in which scalarized training improves only some objectives while the others degrade. We formalize this phenomenon as cross-objective interference and, to our knowledge, conduct the first systematic study of scalarization a...

Yining Lu, Meng Jiang · 0 citations
#machine learning Preprint Open access Oct 2026

Computationally efficient goodness-of-fit tests through kernelized Stein discrepancy

Models with intractable normalizing constants are widely used in statistics and machine learning. Assessing the adequacy of such models poses significant challenges: obtaining samples from the fitted model often requires sophisticated sampling algorithms. Moreover, model fitting sometimes requires iterative numerical o...

Zhihan Huang, Ziang Niu · 0 citations
#machine learning Preprint Open access Oct 2026

Machine learning Majorana topology using unsupervised and supervised learning

In unsupervised learning, the training data for deep learning does not come with any labels, thus forcing the algorithm to discover hidden patterns in the data for discerning useful information. This, in principle, could be a powerful tool in identifying topological order since topology does not always manifest in obvi...

Jacob Taylor, Haining Pan, Sankar Das Sarma · 0 citations
#machine learning Preprint Open access Oct 2026

PyDPF: A Python Package for Differentiable Particle Filtering

State-space models (SSMs) are a widely used tool in time series analysis. In the complex systems that arise from real-world data, it is common to employ particle filtering (PF), an efficient Monte Carlo method for estimating the hidden state corresponding to a sequence of observations. Applying particle filtering requi...

John-Joseph Brady, Benjamin Cox, Yunpeng Li et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Beyond the Semicircle: Free Diffusion Models with Prescribed Equilibria

A growing class of machine-learning objects -- covariance and Gram matrices, kernel and attention matrices, MIMO channel matrices, density operators -- are naturally spectra rather than coordinate vectors. Building a denoising diffusion model for such data by corrupting eigenvalues coordinatewise is not merely elegant:...

Swagatam Das · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Qubit-centric Transformer for Surface Code Decoding

For reliable large-scale quantum computation, quantum error correction (QEC) is essential to protect logical information distributed across multiple physical qubits. Taking advantage of recent advances in deep learning, neural network-based decoders have emerged as a promising approach to improve the reliability of QEC...

Seong-Joon Park, Hee-Youl Kwak, Yongjune Kim · 0 citations
#machine learning Preprint Open access Oct 2026

DRtool: An Interactive Tool for Analyzing High-Dimensional Clusterings

When faced with new data, we often conduct a cluster analysis to obtain a better understanding of the data's structure and the archetypical samples present in the data. However, the increases in data complexity and dimensionality have made this step very tricky. The large proportion of noise in high-dimensional data bl...

Justin Lin, Julia Fukuyama · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models. This bias undermines fairness and reliability in evaluation pipelines, particularly for tasks like preference tuning and model routing....

Dani Roytburg, Matthew Bozoukov, Matthew Nguyen et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.