Skip to content

Category

machine learning

12,041 papers

#machine learning Preprint Open access Oct 2026

Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators

Generative retrieval has emerged as a powerful paradigm for LLM-based recommendation. However, industrial recommender systems often benefit from restricting the output space to a constrained subset of items based on business logic (e.g. enforcing content freshness or product category), which standard autoregressive dec...

Zhengyang Su, Isay Katsman, Yueqi Wang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Fluids You Can Trust: Property-Preserving Operator Learning for Incompressible Flows

We present a novel property-preserving kernel-based operator learning method for incompressible flows governed by the incompressible Navier--Stokes equations. Traditional numerical solvers incur significant computational costs to respect incompressibility. Operator learning offers efficient surrogate models, but curren...

Ramansh Sharma, Matthew Lowery, Houman Owhadi et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Machine learning modularity

Based on a transformer based sequence-to-sequence architecture combined with a dynamic batching algorithm, this work introduces a machine learning framework for automatically simplifying complex expressions involving multiple elliptic Gamma functions, including the $q$-$\theta$ function and the elliptic Gamma function....

Yi Fan, Vishnu Jejjala, Yang Lei · 0 citations
#machine learning Preprint Open access Oct 2026

Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM

Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and computing challenges in deployment. We present a novel low-rank compression framework to address this challenge. First, we upper bound the change of network loss via layer-wise...

Ryan Solgi, Parsa Madinei, Jiayi Tian et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Circuit realization and hardware linearization of monotone operator equilibrium networks

It is shown that the port behavior of a resistor-diode network corresponds to the solution of a ReLU monotone operator equilibrium network (a neural network in the limit of infinite depth), giving a parsimonious construction of a neural network in analog hardware. We furthermore show that the gradient of such a circuit...

Thomas Chaffey · 0 citations
#machine learning Preprint Open access Oct 2026

Persistence Paradox in Dynamic Science: Evidence from the Deep Learning Revolution

Persistence is often regarded as a virtue in science. In this paper, however, we challenge this conventional view by highlighting its contextual nature, particularly how persistence can become a liability during paradigm shifts. We focus on the deep learning revolution catalyzed by AlexNet in 2012. Analyzing the 20-yea...

Honglin Bao, Beichen Lu, Kai Li · 0 citations
#machine learning Preprint Open access Oct 2026

Allocation Stability and Wald Inference under Variance-Aware UCB

Allocation stability is often used to justify Gaussian inference from bandit data, but when is it necessary? In this paper, we address this question for a two-armed, fixed-horizon variance-aware UCB policy with bounded reward distributions that may vary with the horizon. We find a sharp criterion in terms of the reward...

Yingying Fan, Yuxuan Han, Jinchi Lv et al. · 0 citations
#machine learning Preprint Open access Oct 2026

What do Reward Models Memorize?

This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.g., model identity, user sampling strateg...

Ivo Verhoeven, Pushkar Mishra, Ekaterina Shutova · 0 citations
#machine learning Preprint Open access Oct 2026

ECGLight: Compute-Light Framework For Paper ECG Digitization and Myocardial Infarction Screening

Electrocardiography (ECG) is one of the most widely used tests for diagnosing cardiovascular disease. Yet several remote clinics still utilize paper ECG printouts for their analysis due to limited connectivity and computational capacity. As a result, vast numbers of physical ECGs obtained in remote areas still remain i...

Shreyasvi Natraj, Cyrus Achtari, Felice Gragnano et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam

Local sharpness, defined by the largest Hessian eigenvalue $\lambda_1$, sets the maximum stable gradient update size, but its computation would usually require running Lanczos or Hessian-vector products. However, we notice that even a single Armijo backtracking line search already contains this information with just a...

Ashmitha R, J\"org Frochte · 0 citations
#machine learning Preprint Open access Oct 2026

Evaluating Time Series Foundation Models for Electricity Price Forecasting: Contamination Risk, Distributional Shifts, and Covariate Dependence

Time series foundation models (TSFMs) have shown strong zero-shot forecasting performance, but their generalization in covariate-driven, non-stationary settings is underexplored. Electricity price forecasting (EPF) presents a challenging testbed due to complex temporal dependencies, distributional shifts, and strong re...

Zhenghua Pan, Ahmed Aziz Ezzat · 0 citations
#machine learning Preprint Open access Oct 2026

In-Context Residual Calibration for Uncertainty Quantification of Energy Time Series over Graphs

Accurate energy demand forecasting is essential for the reliable operation and planning of modern sustainable energy systems. Spatial-temporal graph neural networks (STGNNs) have recently achieved strong performance in point forecasting by jointly modeling temporal dynamics and relational dependencies across interconne...

Keivan Faghih Niresi, Alice Cicirello, Olga Fink · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.