Skip to content

Category

machine learning

12,457 papers

#machine learning Preprint Open access Oct 2026

When to Intervene? State-Aware Sparse Manipulation in Federated Reinforcement Learning

Federated reinforcement learning (FRL) enables distributed agents to collaboratively train decision-making policies, but its decentralized training process also exposes global policy learning to Byzantine manipulation. Existing poisoning attacks primarily focus on how to construct malicious updates, while trajectory-le...

Shutong Zheng, Sijia Chen · 0 citations
#machine learning Preprint Open access Oct 2026

Rare Gate Disagreements Can Limit Plasticity: When Gradient Flow Mispredicts Finite-Batch SGD

Population gradient flow is a common tool for reasoning about how neural networks adapt, including after pretraining. We show that it can mispredict finite-batch stochastic gradient descent (SGD) qualitatively, and we trace the discrepancy to a specific mechanism. In a two-unit ReLU regression, a source task drives the...

Ruoyu Zhao, Mingxuan Zhang, Jianbo Dai et al. · 0 citations
#machine learning Preprint Open access Oct 2026

MotiveMob: Motivation as Semantic Action for Closed-Loop Human Mobility Generation

Human mobility generation, an important task in urban research, synthesizes trajectory data for urban planning and transportation management. Human mobility can be characterized as a "why-where-when" decision process: people form an intention to move and then determine where and when the corresponding activity will tak...

Mengkun Gao, Zengqing Wu, Renhe Jiang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Policy Alignment: New Signals for Membership Auditing in On-Policy Distillation

On-policy distillation (OPD) trains a student model by aligning its policy with a teacher model on trajectories generated by the student model itself. Through this process, the student policy moves toward the teacher on the prompts used for distillation. However, these prompts are often private and costly, creating a n...

Yilong Yang, Wenzhuo Shang, Yule Liu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

MC-TRCM: Observation-Aware Recursive Fusion for Incomplete Mobile and Wearable Mental-Health Feature Views

Public mobile and wearable mental-health datasets often provide summarized feature tables rather than synchronized raw sensor streams. In these releases, each anchor corresponds to a survey or label time and may combine phone or wearable summaries, prior symptom scores, demographics, clinical variables, and source-avai...

Wentao Wang, Lifeng Han, Zining Ren et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Learning from Hetero Density for Cryo-EM Protein Reconstruction

Reconstructing protein structures from cryo-electron microscopy (cryo-EM) maps is essential for understanding macromolecular assemblies. Although learning-based methods have improved protein reconstruction, information from hetero components remains underused. Our analysis finds both false predictions and reference pro...

Xu Han, Chaozhuo Li, Xiaowei Yuan et al. · 0 citations
#machine learning Preprint Open access Oct 2026

CoPoE: Multimodal Fusion via Decomposable Disease-Coordinate Product-of-Experts for Missing-Modality Alzheimer's Diagnosis

Multimodal Alzheimer's disease (AD) diagnosis benefits from integrating heterogeneous clinical, imaging, genomic, and biomarker evidence, but clinical cohorts frequently suffer from irregular modality missingness. Existing fusion methods often synthesize absent inputs, risking the introduction of artificial surrogates,...

Chihun An, Ikbeom Jang · 0 citations
#machine learning Preprint Open access Oct 2026

From a Prompt to Repertoires: Evolving Functional REpertoires Enable LLM Continual Learning

Continual learning remains challenging for large language models, which must enable models to acquire new skills and knowledge without degrading existing capabilities. Existing approaches typically address this challenge by carefully designing how model parameters are updated. In contrast, prompt optimization avoids co...

Fengyuan Liu, Yue Wang, Hangxi Guo et al. · 0 citations
#machine learning Preprint Open access Oct 2026

SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces

Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces. However, this paradigm inherently suffers from prohibitive inference-time overhead and external dependencies. In this paper, we explore whether an MLLM...

Rongxue Li, Meng Yang, Yiru Mao et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Bernoulli Flow Models: Self-Consistent Generative Modeling for Binary Data

Binary diffusion models typically require a large number of function evaluations (NFEs) to generate high-quality samples, making practical inference computationally expensive. Reducing NFEs while preserving sample quality without distillation or additional training remains a significant challenge. Existing binary diffu...

Hao Mo, Liying Yang, Shumin Yao et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Sample-Efficient Generative Conformal Prediction

Generative conformal prediction builds uncertainty sets from samples of a conditional generator, which are efficient only when the samples represent the response distribution well. This can require many samples, each of which can be costly, as in large diffusion models and scientific simulators, so the sampling budget...

Minxing Zheng, Shixiang Zhu · 0 citations
#machine learning Preprint Open access Oct 2026

Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers

In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component. Existing low-rank PTQ approaches, however, either optimize low-rank compensation and residual quanti...

Shiwen Wang, Pengxiang Zhao, Xiaoming Yuan · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.