Skip to content
Open access

On-Orbit Computing and Caching: A Distributed Regret Learning Approach

2026 · IEEE Open Journal of the Communications Society · Vol 7, pp. 10067-10082 · 0 citations · 41 references
Computer Science

TL;DR

A distributed game-theoretic framework for joint task offloading and service caching in STINs, providing provable equilibrium guarantees is developed, formulated as a non-cooperative stochastic game among IoT devices that autonomously determine their computing, association, and satellite caching strategies.

Abstract

Satellite-Terrestrial Integrated Networks (STINs) leverage the global reach of satellite systems to push onboard computing and caching resources toward the network edge, enabling truly ubiquitous, anytime-anywhere services for remote Internet of Things (IoT) applications. In this context, efficiently orchestrating the constrained onboard computing and caching resources under stochastic service demands in a scalable, low-complexity, and resilient manner remains a critical challenge. Existing solutions primarily rely on centralized optimization or multi-agent learning techniques, which struggle to cope with the non-stationarity arising from the interdependence of autonomous agent decisions. In this work, we take a step further and develop a distributed game-theoretic framework for joint task offloading and service caching in STINs, providing provable equilibrium guarantees. Specifically, the joint problem is formulated as a non-cooperative stochastic game among IoT devices that autonomously determine their computing, association, and satellite caching strategies to minimize their end-to-end latency subject to energy and cache capacity constraints. The formulated game is proven to converge to a Correlated Equilibrium (CE), which generalizes the Nash Equilibrium (NE) to correlated, probabilistic strategy profiles across devices. Two distributed no-regret learning algorithms, operating under different information availability and rationality regimes, are introduced to derive the CE. The effectiveness and efficiency of the two no-regret learning algorithms are validated through extensive simulations, considering alternative equilibria, learning-based methods, baseline computing schemes, and varying network and algorithm configurations.

Read PDF

Similar papers

Joint Optimization of Delay and Energy Efficiency for UAV Task Offloading and Cooperative Scheduling

The growing demand for multimedia services in Internet of Things (IoT) networks has significantly increased the traffic load on backhaul links, making Mobile Edge Caching (MEC) a key technology for reducing content delivery latency. Unmanned Aerial Vehicles (UAVs) can serve as mobile aerial caching nodes that complement fixed ground infrastructure, but their small cache size and limited battery life restrict how long and how effectively they can operate. In addition, current approaches often optimize caching decisions, user association, and flight trajectories separately, without considering their interactions under tight energy constraints. In this paper, we formulate a joint optimization problem that aims to minimize the average content retrieval delay in an energy-constrained multi-UAV cooperative caching system. We then propose a deep reinforcement learning (DRL) framework based on the Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) algorithm, in which each UAV is trained as an independent agent under a centralized-training and decentralized-execution scheme. Simulation results show that our method outperforms several heuristic and non-cooperative reinforcement learning baselines in terms of cache hit rate and energy efficiency. Specifically, the proposed method reduces the system's average content retrieval delay with a maximum reduction of 9.1% and effectively guarantees an average cache hit rate of 62.45%, maintaining a sustained remaining energy margin over baseline methodologies.

Tao Zhang, Tao Xu, Zekai Liu et al. · 0 citations
Open access 2026

Cooperative Task Offloading in Mobile Edge Computing via an Improved MASAC Framework

An adaptive Beta-policy and delayed-update multi-agent soft actor-critic method, abbreviated as ABDMASAC, which uses a Beta policy to model bounded actions and achieves a better overall trade-off than the selected MASAC-backbone and on-policy MARL baselines under the considered simulation settings.

Zheng Yao, Jie Liu, Changjun Deng et al. · 0 citations
Open access Aug 2026

Joint Task Offloading and Resource Allocation with Data Caching in UAV-Aided Mobile Edge Computing Networks for Latency-Sensitive Applications

Simulation results confirm that the proposed JORC framework substantially reduces latency, energy consumption, and overall system cost, while increasing the successful task completion ratio compared to existing baseline approaches.

Tanmay Baidya, Sangman Moh · 0 citations
Open access Jul 2026

MULTI-AGENT REINFORCEMENT LEARNING FOR TASK OFFLOADING AND RESOURCE ALLOCATION IN MEC SYSTEMS

This paper addresses the joint task offloading and resource allocation problem in multi-user MEC systems and proposes a decentralized control framework based on Multi-Agent Reinforcement Learning (MARL), which achieves lower total system cost and faster convergence than the full-local, full-offload, and heuristic baselines.

Youssef Oukissou, Mohamed Amine Meddaoui, Ayoub Belaidi et al. · 0 citations
Open access Jul 2026

TWO-AGENT REINFORCEMENT LEARNING FOR TASK OFFLOADING IN IOT-MEC NETWORKS

The rapid proliferation of Internet of Things (IoT) devices has placed unprecedented pressure on the network edge, where applications such as augmented reality, real-time analytics, and autonomous navigation demand low latency and tight energy budgets that traditional cloud-centric architectures cannot meet. Multi-access Edge Computing (MEC) addresses this gap by relocating computation closer to end users, but the core question of where and how each task should be executed remains open: rulebased and single-objective offloading strategies fail to simultaneously balance service latency, energy efficiency, and user experience under dynamic, large-scale conditions. In this paper we propose TARLOT (Two-Agent Reinforcement Learning Offloading Tasks), a cooperative framework for threetier IoT–MEC–Cloud environments. TARLOT decouples the offloading decision from the resourceallocation problem and assigns each to a dedicated Q-learning agent, so that the two subproblems are specialised independently while still being optimised jointly. The framework is evaluated on PureEdgeSim under heterogeneous IoT workloads, device densities ranging from 200 to 2,400, and diverse application profiles, and is compared against five widely-used baselines (Random, Round-Robin, Trade-Off, Pure-Edge, and Pure-Cloud). At 2,400 devices, TARLOT delivers an average service time of 1.1 s (against 4.3 s for Pure-Cloud), a Quality of Experience of 0.77 (against 0.22 for Pure-Cloud), a task-failure rate below 2 % (against nearly 14 % for Pure-Cloud), and a per-device energy consumption of only 3.6 W (against 11.2 W for Pure-Cloud) — roughly a 68 % reduction. Balanced CPU utilisation across the local, edge, and cloud tiers further confirms that TARLOT prevents resource bottlenecks, establishing it as a practical solution for next-generation large-scale IoT deployments.

Oussama Lagnfdi, Marouane Myyara, A. Darif · 0 citations
2026

Federated Learning of Satellite Aided Computation in LEO Ubiquitous Edge Computing Networks

With the rapid development of artificial intelligence and low-earth orbit (LEO) satellite edge computing technology, there has been a rapid increase in the demand for intelligence networks and services among users in remote areas, such as federated learning (FL). We propose a satellite aided computation FL (SACFL) system in LEO ubiquitous edge computing (UEC) networks, aiming at improving the efficiency of FL tasks in remote areas. In the considered deployment scenario, the following key factors are considered: 1) terrestrial users in remote areas; 2) offloading data to satellites for aided computation; and 3) global aggregation on satellite. However, the highly dynamic characteristics and the uneven distribution of satellite computation resources pose significant challenges to the low-delay requirements of satellite-based FL tasks. To this end, we formulate an optimization problem to minimize delay by jointly considering access selection, computation offloading, aggregation satellite (AgS) selection, and computation resource allocation. To solve the formulated problem, we propose a novel multi-agent alternating (M2A) optimization method. Specifically, three independent agents are trained alternately to make decisions on access selection, computation offloading, and AgS selection. Comprehensive simulations demonstrate that the proposed method outperforms other benchmark algorithms in terms of convergence, delay, and FL accuracy.

Jun-Yi Yang, Yafeng Ma, Zhenyu Xiao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.