The growing use of autonomous artificial intelligence agents in high-risk areas such as -driving cars, medical triage tools, military robots, and disaster-response drones, has made AI ethics an important concern for researchers. Current alignment methods, especially Reinforcement Learning from Human Feedback, have show...
S. Liyanage, S. Rajasingham· Journal of Multidisciplinary...· 0 citations
This repository contains the MATLAB files associated with the paper entitled “Hierarchical Adaptive Sliding Mode Control with Deep Reinforcement Learning for Solar-Assisted Fuel Cell Hybrid Electric Vehicles.” The deposited files support the numerical simulations and performance evaluation reported in the manuscript. T...
Sahbi Boubaker, Khalil Jouili, Mehdhar S. A. M. Al-Gaashani et al.· Zenodo (CERN European Organi...· 0 citations
Abstract This paper proposes a Twin Delayed Deep Deterministic Policy Gradient (TD3)-based adaptive fractional-order proportional-integral-derivative (FOPID) controller for joint load-frequency and voltage regulation in a renewable-rich power system. The proposed controller adaptively tunes five FOPID parameters, namel...
Nitrogen leaching threatens groundwater quality, making production retention and efficient water–fertilizer use joint management priorities. Field trials evaluate a limited set of treatments, while adaptive policies generate long sequences of decisions. We propose counterfactual rollout-constrained offline reinforcemen...
Hao Jin, Xiao Guo, Kelin Hu et al.· AgriEngineering· 0 citations
Large language models increasingly emit confidence reports, predictive distributions, and typed decisions that determine whether a system answers, abstains, retrieves evidence, or spends more computation. We survey calibration-aware reinforcement learning (RL), in which a reported probability is scored by the reward, c...
Abstract The optimization of logic gates plays a crucial role in Electronic Design Automation (EDA), typically managed by applying deterministic, rule-based heuristics to And-Inverter Graph (AIG) models of a circuit. As circuit complexity rises to billions of gates, these heuristic methods find it increasingly difficul...
Abigail Lam· Zenodo (CERN European Organi...· 0 citations
The El-Rakhawi Doctrine on Neuro-Adaptive AI and the Prohibition of Cognitive Exploitation by Dr. M. K. A. El-Rakhawi presents the first legal framework defending human free will against Closed-Loop Neuro-Adaptive Systems (CL-NAS). These systems use Reinforcement Learning to exploit neural vulnerabilities and hijack de...
mohamed kamal arafa el-rakhawi· Zenodo (CERN European Organi...· 0 citations
The El-Rakhawi Doctrine on Neuro-Adaptive AI and the Prohibition of Cognitive Exploitation by Dr. M. K. A. El-Rakhawi presents the first legal framework defending human free will against Closed-Loop Neuro-Adaptive Systems (CL-NAS). These systems use Reinforcement Learning to exploit neural vulnerabilities and hijack de...
mohamed kamal arafa el-rakhawi· Zenodo (CERN European Organi...· 0 citations
Unmanned Aerial Vehicles (UAVs) have emerged recently due to rapid improvements in wireless technology and low-cost equipment, advancement in networking communication techniques, and increased demand from various industries that seek to leverage aerial data to improve their business and operations. As such, UAVs have b...
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026