Skip to content

Category

machine learning

12,457 papers

#artificial intelligence Preprint Open access Oct 2026

Fully Interpretable Minimal Transformers: From Geometry to Algorithm

We present a framework for building and interpreting minimal transformer models. By constraining a transformer's embedding dimension and head size to 2, we enable full two-dimensional visualization of its internal representations. Embeddings, query/key/value transforms, attention outputs, residual streams, and decision...

Raneem Mahajne, Toviah Moldwin · 0 citations
#artificial intelligence Preprint Oct 2026

DisParQ: Self-Supervised Part Concepts for Interpretable Vision Foundation Models

Concept-based vision models represent images through an intermediate layer of human-inspectable concepts, so what a model relies on can be traced to those concepts. However, those models are often limited to fixed categories or depend on language to define their concepts. We introduce DisParQ (Discrete Parts with Quant...

Adam Pardyl, Siddhartha Gairola, Sukrut Rao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Artificial intelligence pathways from weather to climate

Deep learning has made rapid advances in weather forecasting: autoregressive models trained on atmospheric reanalyses now rival dynamical models across nowcasting, medium-range, and subseasonal-to-seasonal lead times, producing well-calibrated ensemble forecasts at reduced cost. We review these advances and consider th...

Tom Beucler, J. David Neelin, Hui Su et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Beyond Policy Support: Interaction Constrained Offline Reinforcement Learning for Autonomous Driving

Offline reinforcement learning enables reward-driven policy improvement from fixed datasets without requiring online exploration, making it particularly attractive in safety-critical domains. A central challenge, however, is distribution shift: policy optimization may favor actions that are weakly supported by the offl...

Mahmoud Selim, Cristina Cipriani, Karl Henrik Johansson · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Unrolled Flow Models for Reasoning

Flow matching enables language generation in few steps, but whether additional integration steps improve reasoning remains unclear. We prove that a flow parameterized by a two-layer Transformer can solve graph reachability, with the required number of integration steps increasing with the target's distance from the roo...

Faissal Izermine, Hanru Bai, Oscar Davis et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Few-Shot Learning for Personalised Automated Pain Assessment

Pain perception varies substantially across individuals, making it difficult for population-based classifiers to generalise across all subjects in a dataset. One way to account for subject variability is to train personalised classifiers. In this work, we evaluate Few-Shot Learning, a sub-area of Meta-Learning, as an a...

Heinke Hihn, Ibrahim Eisawy, Patrick Thiam et al. · 0 citations
#artificial intelligence Preprint Oct 2026

CERO: Where and When to Allocate Rollouts for RL Post-Training

Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts. We instead study how to coordinate a finite rollout budget over the entire training horizon. We formulate this problem using a concave surrogate utility of cumulative prompt exposure and intro...

Yi-Ming Zong, Yi-Ge Wang, Xin-Ting Hu et al. · 0 citations
#artificial intelligence Review Oct 2026

Coding-Agent Benchmarks Should Match Their Users'Task Flows

The evaluation of coding agents generally strives to be as realistic as possible. In our study, we collect 4,782 agent sessions of real software engineers in JetBrains IDEs, which we call Production Sessions. Since our subject is interactive agents, we study the sessions with at least three user messages (33% of the sa...

I. Slinko, Yaroslav Golubev, Sergey Titov · 0 citations
#artificial intelligence Preprint Oct 2026

Collaborative Reasoning Distillation via Cross-Feedback and Coherent Curation

Reasoning capabilities are critical for advancing Large Language Models, yet current approaches either require massive computational budgets or struggle to effectively distill reasoning to smaller models. Standard distillation methods rely on outcome-based rewards, failing to distinguish between sound reasoning and luc...

Tae-hong Kim, Seunggeun Cho, Dong-Su Han · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Align Before You Combine: Reference Space Calibration for Supervision Without Ground Truth

We introduce a calibration-first framework that produces supervision scores without access to ground-truth labels or a shared annotation space. Our framework aligns subset-specific scorers using a synthetic ordinal reference space before fusion. This reference space is constructed from ordered calibration features that...

Jackson Eshbaugh, Jorge Silveyra · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Safe on Average, Unsafe in the Tail: When Is the Episodic-Cost Tail Controllable?

Safe reinforcement learning seeks policies that maximize return while satisfying constraints on cumulative cost. Most methods impose these constraints on expected episodic cost. Consequently, standard evaluations report mean episodic cost without characterizing how cost is distributed across episodes. A policy that sat...

Samuel Tetteh, Cody Fleming · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SpatialUQ: Post-Hoc Uncertainty Quantification from Spatial Consistency in Black-Box Vision Models

Clinical vision models are often deployed as frozen black boxes with no access to internals, retraining, or ground truth at inference time. We introduce \textbf{SpatialUQ}, a post-hoc uncertainty method using only output probabilities. It measures the Jensen-Shannon divergence between the global prediction and the mean...

Md Kawsher Mahbub, Milon Biswas, Mirza Niaz Morshed et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.