Skip to content

Category

reinforcement learning

1,901 papers

A TSP solving model based on multiscale feature fusion and adaptive context awareness

Transformer-based deep reinforcement learning for the Traveling Salesman Problem (TSP) often struggles to capture spatial topology and avoid local optima. To address this, we propose a novel model featuring a Multi-Scale Grid Attention Encoder (MSGAE) to fuse local and global spatial features, alongside a bottleneck-en...

Pingping Dai · 0 citations
#reinforcement learning Open access Oct 2026

Decision-Path Inverse Reconstruction: Observed Decision Outcomes as Structural Constraints on Latent Decision Paths

This working paper proposes Decision-Path Inverse Reconstruction (DPIR), an inverse analytical operation within the Atlas Insight Method (AIM). DPIR treats an observed decision outcome not only as the endpoint of a judgment process, but also as structural information that constrains the set of decision paths capable of...

Miho Osawa · 1 citation
#reinforcement learning Open access Oct 2026

PEMBERDAYAAN GURU DALAM PEMBUATAN MODUL PEMBELAJARAN BERBASIS DEEP LEARNING MELALUI KOLABORASI GURU PADA MGMP BAHASA INDONESIA SE-KOTA PALANGKA RAYA

The implementation of the Deep Learning approach within the Kurikulum Merdeka framework requires teachers to design instructional modules that facilitate meaningful, mindful, and joyful learning experiences. However, a preliminary needs assessment among teachers in the Indonesian Language Subject Teachers' Consultation...

Nuryeni Nuryeni, Hari Windu Asrini · 0 citations
#reinforcement learning Book Open access Oct 2026

ElasticScale: Elastic Large Language Model Reinforcement Learning Training on Heterogeneous Mobile Edge Clusters

ElasticScale is presented, an elastic orchestration system that organizes heterogeneous accelerators into disaggregated rollout and trainer instances, via a HeterogeneousRayWorkerGroup abstraction that manages non-uniform hardware topologies and a multi-instance Federated Weight Averaging protocol that aggregates updat...

Wei-An Lin, M. Reza, Talha Nayyar et al. · 0 citations
#reinforcement learning Open access Oct 2026

Identifiability of Ensemble Disagreement under One-Step Distillation

Model-based reinforcement learning gates imagined transitions by how much an ensemble disagrees, and diffusion world models are distilled into one-step students to make that imagination affordable. Whether distillation preserves the score the gate reads is rarely checked; we show that it need not. Matching the teacher'...

keyush nisar · 0 citations
#reinforcement learning Open access Oct 2026

Gymnasium: A Standard Interface for Reinforcement Learning Environments

Gymnasium is an open-source library providing an API for reinforcement learning environments. Its main contribution is a central abstraction for wide interoperability between benchmark environments and training algorithms. Gymnasium comes with various built-in environments and utilities to simplify researchers' work al...

Mark J. Towers, Ariel Kwiatkowski, Jordan K. Terry et al. · 0 citations
#reinforcement learning Dataset Open access Oct 2026

Database and code for punching shear of FRP-reinforced flat slabs: a leakage-aware benchmark

This dataset accompanies the manuscript "Do machine learning models outperform design codes for the punching shear strength of FRP-reinforced flat slabs? A leakage-aware benchmark with explainable analysis." It contains a database of 87 punching shear tests on interior slab–column connections reinforced with glass or c...

Sreekar Chand Kuruganti, Pavani Taliakula · 0 citations
#reinforcement learning Open access Oct 2026

Manajemen Pemanfaatan Teknologi Digital dalam Mengoptimalkan Kualitas Pembelajaran Akidah Akhlak pada MAN 1 HSU

Digital transformation in Islamic education provides significant opportunities while simultaneously creating challenges in maintaining a balance between technological innovation and character value reinforcement. This study aims to analyze the management of digital technology utilization in optimizing the quality of Ak...

Syahrani Syahrani, Supriadi Supriadi, Winarsih Winarsih et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.