Aug 2026· ACM Transactions on Evolutionary Learning and Optimization· 0 citations· 47 references
TL;DR
BRPG-MAP-Elites is introduced, a new QD-RL algorithm adopting a strictly Markovian actor-critic architecture within the MAP-Elites framework that demonstrates a 43% improvement in average QD-scores over DCRL-MAP-Elites and achieves higher robustness in the generated policies.
Abstract
The family of Quality-Diversity algorithms, such as MAP-Elites, aims to generate a large collection of diverse and highperforming solutions using Evolutionary Computation. Despite their success in domains like evolutionary robotics, relying heavily on random mutations inspired by Genetic Algorithms (GA) makes MAP-Elites inefficient in evolving high-dimensional solutions. This limitation motivated the creation of PGA-MAP-Elites, which incorporated policy gradient (PG) from deep reinforcement learning and enabled evolution on large neural networks. DCRL-MAP-Elites, the latest successor of PGA-MAPElites, further advanced the performance leveraging an actor-critic training that is conditioned on the descriptor. However, the critic evaluation in DCRL-MAP-Elites is made in a non-Markovian manner, which could mislead the neuroevolution with incorrect gradient signals. Therefore, we introduce BRPG-MAP-Elites, a new QD-RL algorithm adopting a strictly Markovian actor-critic architecture within the MAP-Elites framework. BRPG-MAP-Elites learns to predict the impact of actions on both behavior and fitness, using this information to create alternative descriptor-conditioned mutations. Following a comprehensive evaluation on a wide range of locomotion control tasks, our method demonstrates a 43% improvement in average QD-scores over DCRL-MAP-Elites and achieves higher robustness in the generated policies.
Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models have primarily focused on improving optimization stability and training efficiency, while treating the expert selection as a fixed component. Since routing determines the s...
Hong-Yi He, Zheng-Wen Lin, Xiao Liu et al.· 0 citations
This study suggests that RLDMGO can serve as a viable and adaptive solver for complex optimization problems and achieves a competitive ranking among fourteen evaluated state-of-the-art competitors.
This work introduces Elite-Weighted Supervised Fine-tuning (EW-SFT), which uses reward to guide elite selection of high-scoring molecules, and updates the model by its own pretraining loss on that set, and consistently outperforms the corresponding native optimizers.
Shiyun Wa, Yifei Wang, A. G. Green et al.· 0 citations
This work proposes LaRes, a novel hybrid framework that achieves efficient policy learning through reward function search by leveraging large language models to generate the reward function population, guiding RL in policy learning.
Pengyi Li, Hongyao Tang, Jinbin Qiao et al.· Neural Information Processin...· 6 citations
The Narwhal Optimization Algorithm is a recent swarm metaheuristic that, like most population-based optimisers, is prone to premature convergence, is sensitive to random initialisation, and relies on a rigid, schedule-driven exploration–exploitation balance. This paper develops and rigorously evaluates two enhanced v...
A. Al Tawil, S. Z. Hashim, Hanaa Fathi et al.· Scientific Reports· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.