Skip to content

Category

reinforcement learning

1,990 papers

#large language models Open access Oct 2026

VLM-Mapper: Multimodal vision-language models for prior guided lane-level HD map construction

The automated construction of High-Definition (HD) maps from remote sensing data is essential for modern intelligent transportation systems and spatial data infrastructure. While imagery provides a scalable solution for lane-level HD map construction, traditional discriminative models often struggle in complex scenario...

Haofeng Xie, Huiwei Jiang, Yibing Xiong et al. · 0 citations

Pre-control scheme generation for extreme-weather power imbalance risk using multi-resource frequency regulation and deep reinforcement learning

To address the power imbalance risk between renewable energy output and load demand under extreme weather conditions, this paper proposes a pre-control scheme generation method based on the integration of multiple frequency regulation resources and deep reinforcement learning. First, mechanism models for wind power and...

Ze-Xin Mu, Yuan-Ting Hu, Hong-Yu Chen et al. · 0 citations
#reinforcement learning Open access Oct 2026

Machine Learning for Logic Gate Synthesis and Optimization in Electronic Design Automation

The increasing complexity of integrated circuits has made logic synthesis and gate-level optimization important bottlenecks in electronic design automation (EDA). Conventional synthesis flows rely on deterministic transformations, handcrafted heuristics, and repeated evaluation of large design spaces. Machine learning...

Pauleen Racy Lao · 0 citations
#reinforcement learning Open access Oct 2026

Exploration Cost: The Exploration Coefficient in Onboard Reinforcement-Learning Schedulers Is an Independent Driver of Battery Aging

Reinforcement-learning (RL) schedulers are now flying on real spacecraft: NASA’s Carruthers Geocorona Observatory (launched 24 September 2025, per NASA’s own mission status page — not itself a LEO mission; it orbits near the Sun–Earth L1 point) uses deep RL as its default operational scheduler for long-horizon operatio...

John Goodman · 0 citations
#reinforcement learning Open access Oct 2026

CORTEX: An Architecture for Persistent Cognitive Agents. With Verified Persistence, Typed Failure Semantics, and Bounded Cognition

CORTEX: An Architecture for Persistent Cognitive Agents. With Verified Persistence, Typed Failure Semantics, and Bounded Cognition Chloe J. Tully Independent Researcher Engineer https://orcid.org/0009-0007-5661-7332 https://doi.org/10.5281/zenodo.23183421 Tamworth NSW AUSTRALIA October 2026 --- Abstract Long-running co...

Chloe Tully · 0 citations
#reinforcement learning Open access Oct 2026

Machina Mirabilis (GPT-1900)

Machina Mirabilis (GPT-1900) investigates whether a language model trained from scratch on historical text can generate conceptually useful explanations of observations associated with later developments in physics. The project reports a 3.3-billion-parameter transformer and approximately 22 billion tokens of filtered...

Michael Hla · 0 citations
#reinforcement learning Open access Oct 2026

AI-Based Optimization of Logic Gates for Improved Processor Performance

This literature-based review looks at how artificial intelligence, including machine learning and reinforcement learning, can help optimize logic gates and circuits in processor design. It explains how gates, ALUs, and processor performance connect, compares conventional logic synthesis with AI-assisted methods, and su...

John Earl Lizano · 0 citations
#reinforcement learning Open access Oct 2026

Deep reinforcement learning applied to statistical arbitrage investment strategy on cryptomarket

Considerando el aumento al acceso a la información de mercado, en particular el libre acceso a información detallada sobre transacciones de cryptomonedas, junto con la compleja y dinámica propiedad de los mercados financieros, donde se requieren cada vez estrategias de inversión más sofisticadas, el aprendizaje reforza...

Gabriel Vergara Schifferli · 0 citations
#reinforcement learning Open access Oct 2026

Review of: "Protein Structure Prediction in the 3D HP Model Using Deep Reinforcement Learning"

What the work claimsThe authors treat folding in the 3D Hydrophobic-Polar lattice model as a sequential decision problem: the chain is built as a self-avoiding walk on the cubic lattice, one residue per step, and a Deep Q-Network with experience replay and a target network learns where to place the next residue, with t...

Evgeny V. Arsentyev · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.