Skip to content

Category

machine learning

12,457 papers

#artificial intelligence Preprint Open access Oct 2026

Practical Feasibility of Gradient Inversion Attacks in Federated Learning

Gradient inversion attacks are often presented as a serious privacy threat in federated learning, with recent work reporting increasingly strong reconstructions under favorable experimental settings. However, it remains unclear whether such attacks are feasible in modern, performance-optimized systems deployed in pract...

Viktor Valadi, Lucas Beerens, Mattias {\AA}kesson et al. · 0 citations
#machine learning Preprint Open access Oct 2026

AEGIS: Runtime-Guided GPU Collocation for Multi-Tenant Deep Learning Training

Deep learning training commonly runs on shared multi-tenant GPU servers, where exclusive allocation provides isolation but can leave resources underutilized and increase queueing time. Collocation can improve efficiency, but interference-agnostic placement may cause severe slowdowns, while inaccurate memory information...

Ehsan Yousefzadeh-Asl-Miandoab, B\"u\c{s}ra Karatay Demiray, Florina M. Ciorba et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

PuzzleJAX: A Benchmark for Reasoning and Learning

We introduce PuzzleJAX, a GPU-accelerated puzzle game engine and description language designed to support rapid benchmarking of tree search, reinforcement learning, and LLM reasoning abilities. Unlike existing GPU-accelerated learning environments that provide hard-coded implementations of fixed sets of games, PuzzleJA...

Sam Earle, Graham Todd, Yuchen Li et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Improving Mixup Calibration with Wasserstein Distributionally Robust Optimization

In many real-world applications, ensuring the robustness and stability of deep neural networks (DNNs) is crucial, particularly for image classification tasks that encounter various input perturbations. While Mixup-based data augmentation techniques have been widely adopted to enhance the resilience of trained models ag...

Jiaming Hu, Yeping Jin, Debarghya Mukherjee et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Boosting Large Language Models with Mask Fine-Tuning

The large language model (LLM) is typically integrated into the mainstream optimization protocol. However, it remains underexplored whether maintaining the model integrity is \textit{indispensable} for promising performance. In this work, we introduce Mask Fine-Tuning (MFT), a novel LLM fine-tuning paradigm demonstrati...

Mingyuan Zhang, Yue Bai, Huan Wang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Large Language Models for Cryptocurrency Transaction Analysis: A Bitcoin Case Study

Cryptocurrencies are widely used, yet current methods for analyzing transactions often rely on opaque, black-box models. While these models may achieve high performance, their outputs are usually difficult to interpret and adapt, making it challenging to capture nuanced behavioral patterns. Large language models (LLMs)...

Yuchen Lei, Yuexin Xiang, Rafael Dowsley et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Identifiability Analysis of Linear ODE Systems with Hidden Confounders

The identifiability analysis of linear Ordinary Differential Equation (ODE) systems is a necessary prerequisite for making reliable causal inferences about these systems. While identifiability has been well studied in scenarios where the system is fully observable, the conditions for identifiability remain unexplored w...

Yuanyuan Wang, Biwei Huang, Wei Huang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

When Explanations Compete: Policy-Aware Selection Under Uncertainty

Uncertainty-aware explanation methods often produce several alternatives for the same prediction. Selecting among them requires a policy for balancing prediction confidence, uncertainty, and application constraints. This paper presents a framework for applying such policies to a fixed set of generated explanations. Can...

Helena L\"ofstr\"om, Tuwe L\"ofstr\"om, Johan Hallberg Szabadvary · 0 citations
#machine learning Preprint Open access Oct 2026

Data-driven measures of high-frequency trading

Public data do not identify high-frequency trading (HFT), and standard proxies do not separate liquidity-supplying from liquidity-demanding strategies. We overcome this measurement challenge by training machine learning models on proprietary Nasdaq data to map observed HFT activity to public intraday variables. Applyin...

G. Ibikunle, B. Moews, D. Muravyev et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection

This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies which of the classifier's confident decisions can be trusted. This framework, the Latent Diagnostic Taxonomy, consists of (i) constructing a dimensionality-optimized classifi...

Jaturong Kongmanee, Smile Thanapattheerakul · 0 citations
#machine learning Preprint Open access Oct 2026

A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Models

Graph neural networks (GNNs) are routinely employed for spatiotemporal forecasting, yet their performance across widely used benchmark datasets is inconsistent. Here, we perform an audit of dataset properties and baseline models to assess the quality of the benchmarks, and the robustness of the conclusions drawn from t...

Kenneth Martin, Simon Heilig, Asja Fischer et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Systematic Evaluation of TabPFN-TS and Chronos-2 for Zero-Shot Heat Load Forecasting in District Heating Networks

District heating energy hubs require reliable heat load forecasts for efficient operational scheduling. Forecasting models trained on historical data may require retraining as networks evolve. Zero-shot time-series foundation models and in-context forecasting therefore offer a promising alternative: they can adapt at i...

Ben Spoek, Karim K. Ben Hicham, Kai Derzsi et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.