Skip to content

Category

reinforcement learning

1,990 papers

#reinforcement learning Open access Oct 2026

Architectural Exploration of Reinforcement Learning and Graph Neural Networks for Gate-Level Logic Synthesis in Modern EDA

This review-style study examines how reinforcement learning agents and graph neural networks are being integrated into gate-level logic synthesis for electronic design automation. It discusses how GNN-derived structural embeddings support pre-layout estimation of signal probability and switching activity, how RL agents...

Zandro Guinialope · 0 citations
#reinforcement learning Open access Oct 2026

Sensor to pixels: swarm gathering via image-based reinforcement learning

Abstract This study highlights the potential of image-based reinforcement learning methods for addressing swarm-related tasks. In multi-agent reinforcement learning, effective policy learning depends on how agents sense, interpret, and process local inputs. Traditional approaches often rely on handcrafted feature extra...

Y. Koifman, E. Iceland, E. Koifman et al. · 0 citations
#reinforcement learning Open access Oct 2026

CVP-HDDQN: Research software for hierarchical reinforcement learning in gold futures

CVP-HDDQN is the software companion to the manuscript Auction-state representation and expiring advice for hierarchical reinforcement learning in gold futures. Version 1.0.0 is a code-only release containing the original model, feature, learner and execution-ledger modules, selected historical Python implementation and...

Arman Salehi · 0 citations
#reinforcement learning Open access Oct 2026

CVP-HDDQN: Research software for hierarchical reinforcement learning in gold futures

CVP-HDDQN is the software companion to the manuscript Auction-state representation and expiring advice for hierarchical reinforcement learning in gold futures. Version 1.0.0 is a code-only release containing the original model, feature, learner and execution-ledger modules, selected historical Python implementation and...

Arman Salehi · 0 citations
#reinforcement learning Dataset Open access Oct 2026

Code and data for "Reinforcement learning-based PID for nonlinear temperature control with delayed feedback and measurement noise"

Code, data and notebooks accompanying the manuscript "Reinforcement learning-based PID for nonlinear temperature control with delayed feedback and measurement noise" (Y. Sapazhanov, S. Kadyrov; submitted to Mathematical Models in Engineering, Extrica). The study compares six strategies for tuning PID and nonlinear PID...

Yershat Sapazhanov, Shirali Kadyrov · 0 citations
#reinforcement learning Open access Oct 2026

LLM-Driven Verilog Generation and Verification for RISC-V Processor Design

Large language models (LLMs) are increasingly used in Electronic Design Automation (EDA) to write hardware description code. This paper reviews how LLMs generate and verify Verilog for a RISC-V processor datapath, a core topic in Computer Architecture and Organization. The review is built around the AI-driven logic syn...

RENZ HERALD ELAMPARO · 0 citations
#reinforcement learning Open access Oct 2026

MPPT avancés, suiveurs solaires intelligents et électronique photovoltaïque embarquée

Résumé (FR) Ce document, produit avec l'assistance de Gemini 3 Raisonnement, est publié sous licence Apache 2.0. Il constitue une publication défensive volontaire (antériorité) et entre de ce fait dans l'état de la technique dès sa publication en vertu des législations sur les brevets applicables : art. 54(2) CBE (Conv...

Xavier Pillet · 0 citations
#reinforcement learning Open access Oct 2026

Reconceptualizing AI Alignment: Beyond Behavioral Guardrails and Epistemic Simulations

Contemporary AI alignment research is dominated by two paradigms: behavioral conditioning through Reinforcement Learning from Human Feedback (RLHF), and speculative epistemic frameworks such as "Simulation Theology." Both treat alignment as an external constraint imposed on a value-neutral substrate, and both fail for...

Rémi Leroy · 0 citations
#reinforcement learning Open access Oct 2026

Runtime Adaptive Behavioral Direction Control in Autoregressive Transformers via Closed-Loop Activation Telemetry

Large language models encode behavioral constraints directly within their latent representations. Existing approaches to modifying these behaviors rely on parameter-level adjustments (Supervised Fine-Tuning, Reinforcement Learning from Human Feedback, or permanent weight abliteration). These irreversibly alter model we...

Stefan Beierle · 0 citations
#reinforcement learning Open access Oct 2026

AI-Driven Encryption and Resource Optimization in Cloud Security Frameworks

Abstract Cloud-based systems demand secure yet efficient file access mechanisms. Traditional encryption frameworks often introduce latency and resource overhead, limiting scalability. Building on Astillero's (2026) AI-driven logic gate synthesis, this paper explores how Artificial Intelligence (AI) can optimize encrypt...

Diane Rose Cepe · 0 citations
#reinforcement learning Open access Oct 2026

AI-Driven Encryption and Resource Optimization in Cloud Security Frameworks

Abstract Cloud-based systems demand secure yet efficient file access mechanisms. Traditional encryption frameworks often introduce latency and resource overhead, limiting scalability. Building on Astillero's (2026) AI-driven logic gate synthesis, this paper explores how Artificial Intelligence (AI) can optimize encrypt...

Diane Rose Cepe · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.