2026· IEEE Transactions on Network and Service Management· Vol 23, pp. 6336-6350· 0 citations· 35 references
Abstract
Future 6G services will require strict performance guarantees, especially in terms of delay, end-to-end (e2e) across multiple network domains including packet and radio segments. While deterministic transport and slice-based capacity allocation can improve segment-level performance, ensuring e2e Network Service (NS) performance remains challenging as it requires making decisions Near–Real-Time (Near-RT) on a per-service basis, which does not fit well within the typical centralized control and orchestration hierarchy. Multi-agent systems (MAS), where a number of distributed agents collaborate, has demonstrated its capabilities for such Near-RT control. Agents equipped with Deep Reinforcement Learning (DRL) engines autonomously made traffic routing decisions based on e2e telemetry measurements. In this paper, we extend such MAS solutions for NS traffic routing focused on covering several issues that appear under frequent NS reconfiguration, e.g., caused by end device mobility. In addition, we define a lifecycle for NS operation that includes the initial MAS deployment, model reconfiguration during operation, and NS reconfiguration. The proposed lifecycle requires the definition of DRL training and validation procedures to produce models ready to be deployed with guaranteed performance under certain network conditions. In addition, model selection algorithms are defined for the lifecycle scenarios. In case of NS reconfiguration, a procedure for probe testing the actual network conditions is proposed to improve model selection. Evaluation across a meaningful set of network and traffic scenarios shows that the MAS is able to maintain e2e delay guarantees under all the lifecycle scenarios.
This paper presents a comprehensive framework for artificial intelligence (AI)-enabled autonomous network slicing optimization in 6G systems and investigates the application of advanced machine learning paradigms specifically deep reinforcement learning, federated learning, and generative AI to orchestrate dynamic resource provisioning, cross-slice isolation, and proactive SLA (Service Level Agreement) enforcement.
N. P J, Jeeva Jothi· International Journal of Com...· 0 citations
SOVANET+ is presented, an extended scheduling technique that jointly accounts for service criticality, network load, and wireless link quality to allocate resources adaptively across coexisting Vehicle-to-Everything (V2X) services, supporting its viability for next-generation intelligent transportation systems.
Athanasios Kanavos, Gerasimos Papanikolaou-Ntais, A. Kaloxylos· Electronics· 0 citations
(English) Unlike earlier mobile generations, 6G is expected to support a wide range of applications such as immersive communications, remote healthcare, autonomous transportation, and smart cities. These use cases will significantly increase the number of connected devices and impose stringent requirements on bandwidth, latency, reliability, and energy efficiency. As a result, the networks supporting these services will face major challenges in scalability, resource management, and control. In this context, this doctoral thesis investigates the use of Multi-Agent Systems (MAS) as a foundation for next-generation network control. The goal of this thesis is to design and evaluate MAS-based solutions that improve the intelligence, scalability, and energy efficiency of optical networks across both the optical and packet layers.
The first objective addresses the optical layer by investigating centralized and distributed MAS-based approaches for dynamic spectrum control in point-to-multipoint (P2MP) connections. A centralized solution based on traffic prediction and integer linear programming computes optimal allocations under near-real-time constraints, achieving high spectrum utilization but introducing synchronization and scalability limitations. To overcome these issues, distributed architectures are proposed in which transponder agents perform decision-making locally. Three strategies are studied: a mixed-strategy gaming model, a distributed deterministic algorithm, and a multi-agent reinforcement learning (MARL) approach. The MARL solution achieves the best overall performance by anticipating traffic variations and allocating capacity proactively, while distributed methods significantly improve scalability and robustness. Communication efficiency is also studied with the MARL approach allowing for asynchronous operation and reducing inter-agent messaging. Results show that distributed MAS can approach centralized performance while avoiding bottlenecks and single points of failure.
The second objective focuses on the packet layer, where an extended MAS architecture enables end-to-end near-real-time control of network services (NS) through autonomous flow operation. Routing decisions are driven by telemetry and optimized using Deep Reinforcement Learning (DRL) to minimize delay and operational cost, while agents monitor performance and coordinate with the software-defined networking (SDN) controller. The architecture supports the full lifecycle of a NS, including deployment, dynamic reconfiguration, and handover scenarios. A model-selection approach based on offline training and real-time telemetry is proposed, together with an active probe-testing mechanism and long short-term memory (LSTM) based traffic prediction trained online by flow agents. Simulations demonstrate that transferring of trained models between agents enables accurate predictions and knowledge generation allowing for fast reconfiguration decisions while maintaining QoS over the NS.
This MAS architecture provides the foundation for the third objective where experimental results demonstrate reliable QoS maintenance and effective MAS reconfiguration during operation.
In conclusion, this thesis shows that MAS combined with learning-based decision-making, predictive analytics, and distributed control provide a flexible and effective framework for managing future networks. The proposed solutions improve scalability, adaptability, and energy efficiency while maintaining strict performance guarantees, establishing MAS as a key enabler for intelligent and autonomous 6G networks.
(Català) A diferència de les generacions mòbils anteriors, s'espera que el 6G admeti una àmplia gamma d'aplicacions com ara comunicacions immersives, atenció mèdica remota, transport autònom i ciutats intel·ligents. Aquests casos d'ús augmentaran significativament el nombre de dispositius connectats i imposaran requisits estrictes sobre l'amplada de banda, la latència, la fiabilitat i l'eficiència energètica. Com a resultat, les xarxes que donen suport a aquests serveis s'enfrontaran a grans reptes en escalabilitat, gestió de recursos i control. En aquest context, aquesta tesi doctoral investiga l'ús de sistemes multiagent (MAS) com a base per al control de xarxa de nova generació. L'objectiu d'aquesta tesi és dissenyar i avaluar solucions basades en MAS que millorin la intel·ligència, l'escalabilitat i l'eficiència energètica de les xarxes òptiques tant a la capa òptica com a la de paquets.
El primer objectiu aborda la capa òptica investigant enfocaments centralitzats i distribuïts basats en MAS per al control dinàmic de l'espectre en connexions punt a multipunt (P2MP). Una solució centralitzada basada en la predicció de trànsit i la programació lineal entera calcula assignacions òptimes sota restriccions gairebé en temps real, aconseguint una alta utilització de l'espectre però introduint limitacions de sincronització i escalabilitat. Per superar aquests problemes, es proposen arquitectures distribuïdes en què els agents transponedors prenen decisions localment. S'estudien tres estratègies: un model de joc d'estratègia mixta, un algoritme determinista distribuït i un enfocament d'aprenentatge per reforç multiagent (MARL). La solució MARL aconsegueix el millor rendiment general anticipant les variacions del trànsit i assignant la capacitat de manera proactiva, mentre que els mètodes distribuïts milloren significativament l'escalabilitat i la robustesa. També s'estudia l'eficiència de la comunicació amb l'enfocament MARL que permet el funcionament asíncron i redueix la missatgeria interagent. Els resultats mostren que el MAS distribuït pot aproximar-se al rendiment centralitzat evitant els colls d'ampolla i els punts únics de fallada.
El segon objectiu se centra en la capa de paquets, on una arquitectura MAS estesa permet el control de punta a punta en temps gairebé real dels serveis de xarxa (NS) mitjançant el funcionament autònom del flux. Les decisions d'encaminament es controlen mitjançant telemetria i s'optimitzen mitjançant l'aprenentatge per reforç profund (DRL) per minimitzar el retard i el cost operatiu, mentre que els agents supervisen el rendiment i es coordinen amb el controlador de xarxa definida per programari (SDN). L'arquitectura dóna suport al cicle de vida complet d'una xarxa de xarxa (NS), incloent-hi el desplegament, la reconfiguració dinàmica i els escenaris de traspàs. Es proposa un enfocament de selecció de models basat en l'entrenament fora de línia i la telemetria en temps real, juntament amb un mecanisme actiu de proves de sondes i una predicció de trànsit basada en memòria a curt termini (LSTM) entrenada en línia per agents de flux. Les simulacions demostren que la transferència de models entrenats entre agents permet prediccions precises i generació de coneixement que permeten prendre decisions de reconfiguració ràpides mentre es manté la QoS sobre la NS.
Aquesta arquitectura MAS proporciona la base per al tercer objectiu, on els resultats experimentals demostren un manteniment fiable de la QoS i una reconfiguració MAS eficaç durant el funcionament.
En conclusió, aquesta tesi demostra que el MAS combinat amb la presa de decisions basada en l'aprenentatge, l'anàlisi predictiva i el control distribuït proporciona un marc flexible i eficaç per a la gestió de les xarxes futures. Les solucions proposades milloren l'escalabilitat, l'adaptabilitat i l'eficiència energètica, mantenint alhora garanties de rendiment estrictes, establint el MAS com un factor clau per a les xarxes 6G intel·ligents i autònomes.
(Español) A diferencia de las generaciones móviles anteriores, se espera que 6G sea compatible con una amplia gama de aplicaciones, como las comunicaciones inmersivas, la atención médica remota, el transporte autónomo y las ciudades inteligentes. Estos casos de uso aumentarán significativamente el número de dispositivos conectados e impondrán requisitos estrictos de ancho de banda, latencia, fiabilidad y eficiencia energética. Como resultado, las redes que soportan estos servicios se enfrentarán a importantes retos de escalabilidad, gestión de recursos y control. En este contexto, esta tesis doctoral investiga el uso de Sistemas Multiagente (MAS) como base para el control de red de próxima generación. El objetivo de esta tesis es diseñar y evaluar soluciones basadas en MAS que mejoren la inteligencia, la escalabilidad y la eficiencia energética de las redes ópticas en las capas óptica y de paquetes.
El primer objetivo aborda la capa óptica mediante la investigación de enfoques centralizados y distribuidos basados en MAS para el control dinámico del espectro en conexiones punto a multipunto (P2MP). Una solución centralizada basada en la predicción de tráfico y la programación lineal entera calcula asignaciones óptimas con restricciones casi en tiempo real, logrando una alta utilización del espectro, pero introduciendo limitaciones de sincronización y escalabilidad. Para superar estos problemas, se proponen arquitecturas distribuidas en las que los agentes transpondedores toman decisiones localmente. Se estudian tres estrategias: un modelo de juego de estrategia mixta, un algoritmo determinista distribuido y un enfoque de aprendizaje por refuerzo multiagente (MARL). La solución MARL logra el mejor rendimiento general al anticipar las variaciones de tráfico y asignar capacidad de forma proactiva, mientras que los métodos distribuidos mejoran significativamente la escalabilidad y la robustez. También se estudia la eficiencia de la comunicación con el enfoque MARL, que permite la operación asíncrona y reduce la mensajería entre agentes. Los resultados muestran que el MAS distribuido puede aproximarse al rendimiento centralizado, evitando cuellos de botella y puntos únicos de fallo.
El segundo objetivo se centra en la capa de paquetes, donde una arquitectura MAS extendida permite el control integral y casi en tiempo real de los servicios de red (NS) mediante
A constrained optimization model that supports different management goals through alternative objective functions (latency-aware or power-aware) while enforcing operational constraints, including node capacities, slice-specific latency bounds, and explicit limits on VNF migrations/relocations between scheduling periods is proposed.
R. Moreno-Vozmediano, E. Huedo, R. Montero et al.· Journal of Network and Syste...· 0 citations
This paper proposes a multi-agent reinforcement learning (MARL) framework for TSN scheduling, where each TSN queue is modeled as an autonomous agent and the Heterogeneous-Agent Proximal Policy Optimization (HAPPO) algorithm is employed to explicitly model inter-agent dependencies and jointly optimize service delivery across queues.
Marcos Carvalho, Fatih Temiz, Shavbo Salehi et al.· 0 citations
G mobile networks are increasingly using Artificial Intelligence to manage highly dynamic environments characterized by time-varying traffic demands, user mobility, and heterogeneous resources. The dynamic behavior of User Equipment makes timely and accurate control decisions challenging, while distributed data exchange introduces communication overhead and privacy concerns. These challenges call for scalable and communication-efficient learning mechanisms for Radio Access Network (RAN) orchestration. In this paper, we propose DERRIC-FRL, a decentralized Federated Reinforcement Learning framework to orchestrate RAN intelligent controllers. DERRIC-FRL jointly optimizes controller placement and user power allocation through a selective two-level aggregation mechanism, reducing data exchange to only selected orchestrators and controllers while preserving user privacy and improving overall network performance. Specifically, our method significantly reduces total training communication costs by 34% across inter-domain connections, and up to 77% across intra-domain connections, compared to the FedAvg approach. Furthermore, DERRIC-FRL improves user throughput by up to 53% and 61% compared to the DERRIC and FedAvg baselines across a broad range of simulated scenarios.
Elham Hashemi Nezhad, Eric Samikwa, Torsten Braun· IEEE Conference on Network S...· 0 citations