Skip to content

Category

machine learning

12,457 papers

#artificial intelligence Preprint Oct 2026

Shared-Roadmap Generation and Evaluator for Multi-Agent Path Planning Using Heterogeneous Graph Neural Network

Multi-agent path planning (MAPP) in continuous environments often relies on roadmaps to balance safety and search efficiency. However, traditional roadmap generation methods, such as lattice grids or standard sampling-based approaches, frequently face a trade-off between graph density and the likelihood of finding feas...

Brandon Ho, Nikola Rogers, Seung-Kyum Choi · 0 citations
#artificial intelligence Preprint Oct 2026

How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis

As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern. We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated...

M. Karamat, Christian García · 0 citations
#artificial intelligence Review Oct 2026

An Empirical Study of Agent Skills'Downstream Utility

Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance. Existing studies characterize Skill content and evaluate downstream performance, yet provide limited explanations of how utility depends on content, execution configuration, and multi-Sk...

Yu Cheng, De-Hai Zhao, Zhong-Xin Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Route-Verify-Vote: Procedure-Conditioned Self-Consistency for Mixed-Domain Reasoning

Compositional generalization remains challenging when language models must combine familiar reasoning operations in unfamiliar ways. The Scenario-Based Commonsense Reasoning Evaluation (SCoRE) 2026 tests this ability on three mixed domains absent from training and requires models to identify the complete set of correct...

Xinchen Xiao · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Accelerating Floating-Point Satisfiability Solving via Gradient Normalization

Satisfiability Modulo Theories (SMT) solvers are foundational to software verification, program analysis, and compiler testing, particularly over the theory of Quantifier-Free Floating-Point (QF_FP). While recent optimization-based SMT solvers have successfully applied gradient descent to continuous relaxations of logi...

Yuanzhuo Zhang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Do AI weather models miss extremes?

AI weather models are often reported to underestimate extremes, but most evidence concerns deterministic regression models verified against reanalysis. We evaluate twelve physical and AI forecast models against ECMWF IFS using ten months of European station observations. The evaluation covers 10 m wind, 2 m temperature...

Marvin Vincent Gabler, Roberto Molinaro, Niall Siegenheim et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers

Muon-trained modular-arithmetic transformers can lose accuracy while retaining linearly decodable task information. Adjacent swaps localize five captured unnormalized failures to AdamW readout updates. Multiplying the actual readout displacement by the large feature mean produces a class-dependent logit offset shared a...

Ali Janati, Kaoutar El Maghraoui, Anass Belfatmi · 0 citations
#machine learning Preprint Open access Oct 2026

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift

We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. Our key intuition is that by training the model to distinguish good and bad tokens in a response, we naturally guide the model towards generating good tokens, whil...

Zitong Huang, Gustavo Lucas Carvalho, Deqing Fu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SAE++: Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs

Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their internal visual representations remain difficult to interpret. Sparse Autoencoders (SAEs) provide a scalable way to decompose dense model activations into sparse, interpretable features. However, existing S...

Yusong Zhao, Hengyi Wang, Tanuja Ganu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Enhancing Spectral Embedding through Robust and Flexible Knowledge Transfer in Electronic Health Records

We propose a spectral-based, unsupervised representation learning framework to derive low-dimensional embeddings for clinical concepts and patients in rare disease cohorts from electronic health records, where data are high-dimensional but sample sizes are limited. To overcome this challenge, we incorporate a knowledge...

Feiqing Huang, Zongqi Xia, Rong Ma et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces

Reasoning-trained language models can perform, zero-shot, multi-label tasks that require selecting a small set of relevant labels from a universe of thousands to hundreds of thousands of candidates. We ask how they do it mechanistically, and whether the mechanism can be distilled. We make the question measurable by tre...

Debjyoti Saha Roy · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.