Skip to content

Category

reinforcement learning

1,990 papers

#large language models Book Open access Oct 2026

ns3-GenAI: Integrating Large Language Models with ns3 for AI-Native Network Simulations

Existing ns3/AI bridges target numeric reinforcement learning (RL) pipelines and cannot handle the text-centric prompt/response exchange, structured output validation, and multi-node orchestration that large language models (LLMs) require. We introduce ns3-GenAI, an open framework that augments the ns3 shared-memory in...

Su-Bin Han, Junkyu Hong, Sangheon Pack · 0 citations
#reinforcement learning Open access Oct 2026

Identifiability of Ensemble Disagreement under One-Step Distillation

Model-based reinforcement learning gates imagined transitions by how much an ensemble disagrees, and diffusion world models are distilled into one-step students to make that imagination affordable. Whether distillation preserves the score the gate reads is rarely checked; we show that it need not. Matching the teacher'...

keyush nisar · 0 citations
#reinforcement learning Open access Oct 2026

Gymnasium: A Standard Interface for Reinforcement Learning Environments

Gymnasium is an open-source library providing an API for reinforcement learning environments. Its main contribution is a central abstraction for wide interoperability between benchmark environments and training algorithms. Gymnasium comes with various built-in environments and utilities to simplify researchers' work al...

Mark J. Towers, Ariel Kwiatkowski, Jordan K. Terry et al. · 0 citations
#reinforcement learning Dataset Open access Oct 2026

Database and code for punching shear of FRP-reinforced flat slabs: a leakage-aware benchmark

This dataset accompanies the manuscript "Do machine learning models outperform design codes for the punching shear strength of FRP-reinforced flat slabs? A leakage-aware benchmark with explainable analysis." It contains a database of 87 punching shear tests on interior slab–column connections reinforced with glass or c...

Sreekar Chand Kuruganti, Pavani Taliakula · 0 citations
#reinforcement learning Open access Oct 2026

Manajemen Pemanfaatan Teknologi Digital dalam Mengoptimalkan Kualitas Pembelajaran Akidah Akhlak pada MAN 1 HSU

Digital transformation in Islamic education provides significant opportunities while simultaneously creating challenges in maintaining a balance between technological innovation and character value reinforcement. This study aims to analyze the management of digital technology utilization in optimizing the quality of Ak...

Syahrani Syahrani, Supriadi Supriadi, Winarsih Winarsih et al. · 0 citations
#reinforcement learning Dataset Open access Oct 2026

Database and code for punching shear of FRP-reinforced flat slabs: a leakage-aware benchmark

This dataset accompanies the manuscript "Do machine learning models outperform design codes for the punching shear strength of FRP-reinforced flat slabs? A leakage-aware benchmark with explainable analysis." It contains a database of 87 punching shear tests on interior slab–column connections reinforced with glass or c...

Sreekar Chand Kuruganti, Pavani Taliakula · 0 citations
#reinforcement learning Dataset Open access Oct 2026

Control-Guided Reinforcement Learning for Cooperative Energy Management

Dataset associated with the open access publication "Control-Guided Reinforcement Learning for Cooperative Energy Management" by Isabela Fons Moreno-Palancas, Rubén Ruiz Femenia, Raquel Salcedo Díaz, José A. Caballero, Chanona A del R. (Systems and Control Transactions. 2026, 6, 1558-1564). The dataset includes results...

José Antonio Caballero · 0 citations
#reinforcement learning Open access Oct 2026

Identifiability of Ensemble Disagreement under One-Step Distillation

Model-based reinforcement learning gates imagined transitions by how much an ensemble disagrees, and diffusion world models are distilled into one-step students to make that imagination affordable. Whether distillation preserves the score the gate reads is rarely checked; we show that it need not. Matching the teacher'...

keyush nisar · 0 citations
#reinforcement learning Open access Oct 2026

Training-Data Axes in Imitation Learning Cascade into Hybrid Reinforcement-Learning Fine-Tuning: A Leave-One-Out Bridge Study Under Matched and Cascade Evaluation

Hybrid imitation-learning-to-reinforcement-learning (IL→RL) driving stacks are typically evaluated against a single fixed IL prior, leaving open whether IL training-data quality determines downstream RL outcomes and whether hybrid actuator decoupling (IL steers, RL controls only speed) isolates the speed controller fro...

Laurențiu Carabulea, Claudiu Radu Pozna · 0 citations
#reinforcement learning Open access Oct 2026

EG-CT-MADDPG: Source Code, Trained Models, and Validation Data for Dynamic Truck–Drone Delivery Scheduling

Source code, trained model checkpoints, input instances, and reference results for a dynamic truck–drone delivery scheduling method based on multi-agent deep reinforcement learning. The package accompanies the manuscript and supports reproduction of the reported experiments.

Xinyi Li · 0 citations

Hierarchical collaborative multi-agent reinforcement learning for hot-rolling production planning with order-splitting flexibility

In mass personalized hot rolling, intricate constraints cause load imbalances and low order fulfilment. While order splitting alleviates these bottlenecks, it increases changeover frequency and planning complexity. We propose a bi-level model: the upper level optimizes production cost and time, while the lower minimize...

Ruilin Pan, Xinyu Jin, Jianhua Cao et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.