Skip to content

Category

machine learning

12,457 papers

#artificial intelligence Preprint Open access Oct 2026

\$OneMillion-Bench: How Far are Language Agents from Human Experts?

As language models (LMs) evolve from chat assistants to long-horizon agents capable of multi-step reasoning and tool use, existing benchmarks remain largely confined to structured or exam-style tasks that fall short of real-world professional demands. To this end, we introduce \$OneMillion-Bench (\$OMB), a benchmark of...

Yang Liu, Jiaqi Li, Jun Bai et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Trust, Don't Trust, or Flip: Robust Preference-Based Reinforcement Learning with Multi-Expert Feedback

Preference-based reinforcement learning (PBRL) offers a promising alternative to explicit reward engineering by learning from pairwise trajectory comparisons. However, real-world preference data often comes from heterogeneous annotators with varying reliability; some accurate, some noisy, and some systematically advers...

Seyed Amir Hosseini, Maryam Abdolali, Amirhosein Tavakkoli et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Explanation Multiplicity in SHAP: Characterization and Assessment

SHAP explanations are widely used in high-stakes settings to justify decisions, yet they can differ substantially across repeated runs, even when the model, the input instance, and the prediction are held fixed. Prior work has documented disagreement between explanation methods; we show that substantial disagreement ar...

Hyunseung Hwang, Seungeun Lee, Lucas Rosenblatt et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing

Reliable estimation of feature contributions in machine learning models is essential for transparency, algorithmic fairness, and regulatory compliance. While permutation feature importance is widely used, classical implementations rely on repeated Monte Carlo shuffling, introducing significant computational overhead an...

Albert Dorador · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Towards Scalable Meta-Learning of near-optimal Interpretable Models via Synthetic Model Generations

Decision trees are widely used in high-stakes fields like finance and healthcare due to their interpretability. This work introduces an efficient, scalable method for generating synthetic pre-training data to enable meta-learning of decision trees. Our approach samples near-optimal decision trees synthetically, creatin...

Kyaw Hpone Myint, Zhe Wu, Alexandre G. R. Day et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

CAF\'E: Causal Black-Box Testing of Machine Unlearning

Machine learning models are increasingly deployed as software components that must evolve as requirements change. When specific training records or features must no longer influence a deployed model, machine unlearning aims to remove that influence without retraining from scratch. Because unlearning is often approximat...

Anna Mazhar, Sainyam Galhotra · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping

Connecting pretrained vision and language models usually involves projecting image features into the language model's input space. Inverse-LLaVA reverses this mapping within decoder attention: language states are projected to the visual feature dimension, and modality-specific maps produce residual query, key, and valu...

Xuhui Zhan, Tyler Derr · 0 citations
#artificial intelligence Preprint Open access Oct 2026

An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems

Frontier large language models (LLMs) now reach near-ceiling accuracy on standard mathematical-reasoning benchmarks and gold-medal-level performance at the International Mathematical Olympiad. As these benchmarks saturate and their items leak into training data, a high score no longer shows whether a model reasons robu...

Yuren Hao, Xiang Wan, ChengXiang Zhai · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Heterogeneous-Modal Unsupervised Domain Adaptation via Latent Space Bridging

Unsupervised domain adaptation (UDA) effectively bridges the domain gap between a labeled source domain and an unlabeled target domain, but assumes that the two domains share the same modality. Heterogeneous domain adaptation (HDA) instead handles different feature spaces across domains, yet requires labeled target sam...

Jiawen Yang, Shuhao Chen, Shengtao Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Towards Reasonable Concept Bottleneck Models

We propose a novel, flexible, and efficient framework for designing Concept Bottleneck Models (CBMs) that enables practitioners to explicitly encode and extend their prior knowledge and beliefs about the concept-concept ($C-C$) and concept-task ($C \to Y$) relationships within the model's reasoning when making predicti...

Nektarios Kalampalikis, Kavya Gupta, Georgi Vitanov et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

BEAT: Balanced Frequency Adaptive Tuning for Long-Term Time-Series Forecasting

Long-term time-series forecasting supports a wide range of applications, including weather prediction and electricity demand planning. Frequency-domain methods address this task by decomposing observations into components that describe temporal variations at different scales. However, separate representations do not by...

Zhixuan Li, Naipeng Chen, Seonghwa Choi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Policy Learning with a Language Bottleneck

Modern AI systems such as self-driving cars and game-playing agents can achieve superhuman performance, but often lack human-like generalization, interpretability, and inter-operability with human users. Inspired by the rich interactions between language and decision-making in humans, we introduce Policy Learning with...

Megha Srivastava, Cedric Colas, Dorsa Sadigh et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.