Skip to content

Category

machine learning

12,457 papers

#artificial intelligence Preprint Open access Oct 2026

Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents

Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment...

Wenyu Huang, Xinyu Hou, Pavlos Vougiouklis et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ExperienceIndex: Artifact-Grounded Memory

Knowledge-intensive tasks require answering many questions by reasoning about a shared corpus of artifacts (e.g., court cases, or scientific literature). As humans interact with these corpora, they naturally accumulate experiential knowledge about artifacts, enabling them to quickly identify the complete set of relevan...

Peter Baile Chen, Geoffrey X. Yu, Xinming Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Activation-Aware Weight Tensorization: A Calibration-Time Preconditioner for Tensor-Network LLM Compression

Post-training tensor-network compression replaces Transformer linear layers with Tensor Train (TT) or Tree Tensor Network (TTN) operators, but standard decompositions minimize weight-space Frobenius error rather than functional error under the layer's activation distribution. We propose Activation-aware Weight Tensoriz...

Alessandro Beatini, Marco Maronese, Emanuele Rodol\`a · 0 citations
#artificial intelligence Preprint Open access Oct 2026

TRACK: Telemetry-Based Racing Analysis and Coaching Kit in Sim Racing Games

This paper presents TRACK (Telemetry-Based Racing Analysis and Coaching Kit), which is a framework for analyzing driving performance in sim racing and profiling how individual drivers behave behind the wheel. We report this framework together with its limitations: we calibrate each clustering result against a null, and...

Efe \c{C}ang{\i}r{\i}l{\i}, Murat Kurt · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Beyond Reward Suppression: Near-Optimal Offline Attacks on Warm-Start Bandits with Bounded Rewards

Adversarial attacks on bandits aim to mislead a learner toward a target arm while keeping the attack cost small. Existing attacks typically achieve this by suppressing non-target arms. In practice, however, manipulation such as fake reviews often directly promotes the target item. We study this gap through bounded offl...

Qirun Zeng, Manhin Poon, Xiangxiang Dai et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Temporal Predictive Multiplicity: Equally Accurate Time Series Models Yield Different Forecast Trajectories

Models with near-identical predictive performance can yield substantially different predictions, a phenomenon known as predictive multiplicity. Prior work has mostly studied this at the level of individual scalar outputs. In time-series forecasting, however, predictions across horizons jointly define a trajectory, and...

Emanuele Albini, Francesca Toni, Saumitra Mishra et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Efficient Patch-Based Anomaly Detection Fused with Diffusion Driven Generative Modeling for Semiconductor Wafer Bin Map Open Set Anomaly Detection

Spatial defect signatures on wafer bin maps (WBMs) trace yield loss to specific process faults, yet supervised classifiers recognize only the defect types seen during training, and one-class detectors built on a single mechanism tend to capture either local structural deviations or global distributional violations, but...

Limon Bin Hossain, Md Sadib Rahman Ananta · 0 citations
#artificial intelligence Preprint Open access Oct 2026

KGATE : a Knowledge Graph Embedding Training Environment

Knowledge graph embedding (KGE) models encode the entities and relations of a knowledge graph into a low-dimensional latent space, enabling tasks such as classification or link prediction. Most KGE models follow an autoencoder architecture, in which an encoder projects the knowledge graph into the latent space and a de...

Benjamin Loire, Galadriel Bri\`ere, C\'elia Brahimi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RollVerify: Bridging Efficiency and Accuracy in Long-Tail Rollout Reinforcement Learning

Reinforcement learning is crucial for improving large language models' reasoning and generalization. It relies on massive rollouts whose lengths become increasingly long-tailed as context windows grow. In on-policy training, these long-tail rollouts can result in GPU bubbles, reducing system utilization and limiting RL...

Yongqiang Yao, Jinru Tan, Kaihuan Liang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

An AI-assisted conditioning and geological interpretation workflow for usage in implicit geological modeling

Implicit modeling and Relative Geologic Time are geological modeling techniques that enable more efficient, faster, less biased and more reproducible modeling results. For optimal operation, these techniques require many well-constrained input data. In the framework of the Horizon Europe GO-Forward and MOOI WarmingUP G...

Stefan Carpentier, Jan Diederik van Wees, Eva de Boever et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Training Advisors for LLM Agents from Task Outcomes

Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment. Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution. We introduce Caddie, a method for training critics to provide natur...

Sergei Polezhaev, Barys Liskavets, O. Press et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Stream-Based Active Learning with Cooperative Neural Networks for Data-Efficient Partial Inverse Design: An Automotive Glass Run Channel Case Study

Inverse design in engineering often runs into a simple problem. Each labeled training sample must be produced through expensive simulation, so building a large dataset is slow and costly. This study addresses that problem for partial inverse design, where only some design variables are specified and the rest must be in...

Agung Nugraha, Hyerin Kwon, Heungjun Im et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.