Skip to content

Soft Actor-Critic based regional traffic signal control in connected environment and its application in priority signal control

Sep 2026 · Journal of Intelligent Transportation Systems · Vol 30, pp. 1020 - 1038 · 4 citations · 69 references

TL;DR

A distributed TSC model based on the Soft Actor-Critic (SAC) reinforcement learning algorithm that demonstrates the model’s effectiveness, adaptability, and potential for deployment in intelligent traffic management systems is proposed.

View source

Similar papers

Open access Aug 2026

Deep reinforcement learning-based adaptive traffic signal control in urban networks

This paper introduces a traffic signal control system using deep reinforcement learning to solve the congestion problems at signalized intersections under dynamic traffic conditions. The proposed framework is simulated with the help of MATLAB-SUMO co-simulation framework, the traffic signal control is modeled as a Markov Decision Process (MDP). State space includes traffic density, the queue length, vehicles waiting time, and the actual signal phase whereas the action space comprises of the possible selections of the signal phase. A Deep Q-Network (DQN) is utilized to estimate the optimal state-action value function so that green times can be dynamically allocated based on the changing traffic demand. Multi-objective reward functionality is based on the combined minimization of vehicle delay, queue length, and waiting time and maximization of traffic throughput. Experience replay and target network updates are used to stabilize the learning process. Simulation experiments are conducted in low, medium, and high traffic demand conditions to compare the work of the suggested framework with fixed-time control, actuated control, and classical tabular Q-learning methods. Experimental results demonstrate that the proposed framework reduces average delay by up to 33.9%, decreases queue length by 47.1%, and increases throughput by 25.5% compared to fixed-time control under high-demand conditions. Finally, scalability studies involving networks of up to 16 isolated signalized intersections were conducted to assess computational feasibility and robustness even when the size of the network grows. Comprehensively, the results prove the usefulness of deep reinforcement learning in the creation of intelligent and adaptive traffic signal control services in the city.

Manisha Aeri, K. Purohit, Lata Nautiyal et al. · 0 citations
Open access Aug 2026

Deep reinforcement learning-based traffic signal control in multi-intersection environments: a comparative study of DQN variants

The findings demonstrate the potential of DRL-based traffic signal control in controlled simulation conditions and highlight that algorithm performance is strongly influenced by traffic policy design and environmental complexity.

D. Prastiyanto, A. A. Manaf, Muhammad Ahnaf Maulana et al. · 0 citations
Open access Jul 2026

Dynamic traffic signal scheduling system based on adaptive quad agent Double Deep Q -network algorithm

Simulation results indicate that the proposed Adaptive Quad-Agent Double Deep Q-Network model effectively supports dynamic signal phase adaptation, minimizes congestion, and provides more accurate queue length estimations under complex traffic conditions.

Bharathi Ramesh Kumar, Sachin Salunkhe, S. Shinde et al. · 0 citations
Open access Jul 2026

Adaptive Dynamic Model-Free Neuro-Fuzzy System for Traffic Signal Control

With the development of the economy, urban traffic congestion has become increasingly serious, causing a series of problems such as safety and environmental pollution. To address the challenge of urban traffic congestion problems, traffic signal configuration optimization based on time-varying traffic states and real-time performance is an important direction with practical engineering significance. In this study, we design a traffic signal controller that learns online without requiring a traffic model. A model-free action-dependent adaptive dynamic programming (ADP) that employs two neural networks (an action network and a critic network) to approximate the Hamilton–Jacobi–Bellman equation is employed to provide self-learning optimization ability with varying traffic states. The proposed controller employs a neuro-fuzzy system that functions as an action network to generate control decisions utilizing expertise and mitigates the stochastic exploration inefficiency inherent in conventional ADP. ADP provides reinforcement signals that indicate a reward or punishment for the neuro-fuzzy system to guide it in adjusting its parameters. Subsequently, actions associated with lower cumulative delay are reinforced. The proposed traffic signal controller can reduce the blindness of learning by using experience and the inaccuracy of the traffic model. An artificial bee colony algorithm was employed to train the ADP to meet the real-time requirement for traffic signal control. The controller can adapt to fluctuating traffic states by training continuously, and achieves a smaller average delay in the long run. The simulation results demonstrate that proposed controller achieves a reduced delay through supervised learning with an accelerated training speed.

Tao Li, D. Qian, Xinlan Guo · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.