Skip to content

Category

machine learning

12,457 papers

#artificial intelligence Preprint Open access Oct 2026

Aligning LLMs with Biomedical Knowledge using Balanced Fine-Tuning

Engineering LLMs to accelerate life sciences research requires a robust alignment with biomedical knowledge. We observe that biomedical text exhibits a fundamentally different uncertainty structure from general text: dense low-confidence runs encode epistemic knowledge gaps (dense causal chains, rare entities) rather t...

Zhenchao Tang, Fang Wang, Haohuai He et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization

Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-policy training pipelines, these methods can exhibit unstable optimization dynamics and are prone to perfor- mance collapse. Through empirical analysis, we iden...

Chenliang Li, Adel Elmahdy, Alex Boyd et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation

Modern language models fail a fundamental requirement of trustworthy intelligence: knowing when not to answer. Despite achieving impressive accuracy on benchmarks, these models produce confident hallucinations, even when wrong answers carry catastrophic consequences. Our evaluations on GSM8K, MedQA and GPQA show fronti...

Mohamad Amin Mohamadi, Tianhao Wang, Zhiyuan Li · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Quadratic Direct Forecast for Training Multi-Step Time-Series Forecast Models

The design of learning objectives is central to training time-series forecasting models. Existing learning objectives such as mean squared error mostly treat each future step as an independent, equally weighted task, which leads to the following two challenges: (1) they overlook the label autocorrelation effect among f...

Hao Wang, Licheng Pan, Yuan Lu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Time-multiplexed layer reuse for physical neural networks

Physical neural networks (PNNs) are promising candidates for next-generation computing, but existing demonstrations remain several orders of magnitude smaller than modern digital neural networks, whose recent advances have been driven by rapid growth in trainable parameters. This situation resembles the constraints of...

Kohei Tsuchiyama, Andre Roehm, Takatomo Mihana et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DistDF: Time-Series Forecasting Needs Joint-Distribution Wasserstein Alignment

Training time-series forecasting models requires aligning the conditional distribution of model forecasts with that of the label sequence. The standard direct forecast (DF) approach resorts to minimizing the conditional negative log-likelihood, typically estimated by the mean squared error. However, this estimation pro...

Hao Wang, Licheng Pan, Yuan Lu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Predicting kernel regression learning curves from only raw data statistics

We study kernel regression with common rotation-invariant kernels on real datasets including CIFAR-5m, SVHN, and ImageNet. We give a theoretical framework that predicts learning curves (test risk vs. sample size) from only two measurements: the empirical data covariance matrix and an empirical polynomial decomposition...

Dhruva Karkada, Joseph Turnbull, Yuxi Liu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

HARL-A: An Extensible Benchmark Framework for Heterogeneous Multi-Agent Adversarial Reinforcement Learning in IsaacLab

Progress in adversarial multi-agent reinforcement learning (MARL) for robotics has been hampered by a lack of shared, extensible infrastructure that supports heterogeneous agent morphologies in high-fidelity physics simulation. Existing frameworks either focus on cooperative tasks, rely on simplified physics engines, o...

Isaac Peterson, Christopher Allred, Jacob Morrey et al. · 0 citations
#machine learning Preprint Open access Oct 2026

PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning

Tasks on complex systems require high-precision numerical computation to support decisions. However, current large language models (LLMs), even with enhanced reasoning capabilities, cannot integrate such computations as an intrinsic and interpretable capability with existing architectures. To this end, we propose Physi...

Jingyuan Fan, Purui Liu, Hengbo Xiao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Cross-Modality Controlled Molecule Generation with Diffusion Language Model

The increasing variety of molecular data creates a need for generative models that can flexibly incorporate heterogeneous constraints across modalities. However, existing SMILES-based diffusion models are typically designed for a fixed conditioning modality, and introducing new constraints often requires retraining the...

Yunzhe Zhang, Yifei Wang, Khanh Vinh Nguyen et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Direct Regret Optimization in Bayesian Optimization

Bayesian optimization (BO) is a powerful paradigm for optimizing expensive black-box functions. Traditional BO methods typically rely on separate hand-crafted acquisition functions and surrogate models for the underlying function, and often operate in a myopic manner. In this paper, we propose a novel direct regret opt...

Fengxue Zhang, Yuxin Chen · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Time-o1: Time-Series Forecasting Needs Transformed Label Alignment

Training time-series forecasting models poses unique challenges in loss function design. Most existing approaches adopt temporal mean squared error, but this study reveals two critical limitations: (1) it ignores the presence of label autocorrelation, which biases it from the true label sequence likelihood; (2) it invo...

Hao Wang, Licheng Pan, Zhichao Chen et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.