Skip to content
Open access

LLM-Driven Context-Aware Health Monitoring for Resource-Constrained Edge Devices

Aug 2026 · Electronics · Vol 15, pp. 3579 · 0 citations · 17 references

TL;DR

An LLM-driven, context-aware framework that integrates real-time system metrics, historical data, and task-specific importance levels for anomaly detection and prediction is proposed, enabling proactive intervention before critical operating conditions are reached.

Abstract

Resource-constrained edge devices require efficient and adaptive health monitoring to ensure reliable operation under dynamic workloads. This paper proposes an LLM-driven, context-aware framework that integrates real-time system metrics, historical data, and task-specific importance levels for anomaly detection and prediction. Specifically, the framework forecasts the semantic health state of the edge device five minutes ahead based on recent monitoring observations, enabling proactive intervention before critical operating conditions are reached. Multidimensional metrics, including CPU, memory, temperature, network load, and process information, are transformed into structured time-series representations and used as input to pre-trained machine learning models. A large language model (LLM) acts as an orchestration layer, dynamically selecting the most appropriate predictive model based on system context and resource constraints. This enables adaptive trade-offs between accuracy, latency, and computational cost. Experimental results on Raspberry Pi devices show that the proposed approach achieves comparable or improved performance while reducing resource usage compared to static methods.

Read PDF

Similar papers

Preprint Aug 2026

LLM-Powered Predictive Decision-Making for Sustainable Data Center Operations

This work introduces a novel LLM-based predictive scheduling system designed to enhance operational efficiency while reducing the environmental impact of data centers, using an LLM to predict key metrics such as execution time and energy consumption from source code.

Hanzhao Wang, Jingxuan Wu, Yumeng Li et al. · 0 citations
Open access Aug 2026

Optimizing Latency and Energy Efficiency in Edge-Native Large Language Models (LLMs) for Autonomous Mobile Agents

This study introduces an edge-native framework for optimizing latency and energy efficiency in LLM-enabled autonomous mobile agents and shows decreased communication overhead, increased operational continuity, and faster response times without significantly lowering language comprehension or decision-making precision.

A. Rajalakshmi, D. Saveetha, S. V. Manikanthan et al. · 0 citations
Open access Aug 2026

Edge-cloud collaboration-driven predictive modeling for high-performance computing centers.

Accurate coolant flow prediction is critical for active thermal management in high-performance computing (HPC) centers, yet it is inherently challenged by mixed-timescale dynamics and high-frequency workload surges. Existing deep learning methods often prioritize global accuracy on smoothed stationary trends, which may lead to phase delays during abrupt thermal transients. In addition, high-capacity architectures can introduce non-negligible computational overhead for latency-sensitive, resource-constrained edge controllers. To overcome these limitations, this study proposes a deployment-oriented edge-cloud collaboration (ECC) framework integrated with a transient-aware predictive architecture, named FS-Attention, designed to balance transient responsiveness, engineering deployability, and decision transparency. FS-Attention couples local feature synthesis, temporal-memory encoding, and attention-based temporal refinement to improve coolant-flow tracking under non-stationary operating conditions. Evaluations on the real-world Frontier supercomputer dataset show that the feature synthesis attention (FS-Attention) model achieves competitive full-year prediction accuracy, with a coefficient of determination (R²) of 0.8744 and a root mean square error (RMSE) of 0.0350. Under isolated critical thermal events (CTEs), FS-Attention obtains the lowest RMSE of 0.0757, slightly lower than the Temporal Fusion Transformer (TFT) and 5.61% lower than the standard Transformer. Platform-based profiling further shows an inference latency of 0.016 ms and a parameter size of 323.1 K, suggesting model-side compatibility with facility-side edge execution, while attention-shift analysis provides diagnostic evidence of temporally adaptive model behavior under dynamic thermal conditions.

Shuaiyin Ma, Ye-Ye Cao, Yang Liu et al. · 0 citations
#edge computing Open access Sep 2026

Clinical risk-aware reinforcement learning for latency-constrained healthcare IoT scheduling

The rapid growth of the Healthcare Internet of Things (HIoT) has enabled continuous, real-time patient monitoring through wearable and bedside devices. These systems generate time-sensitive physiological data that are essential for the early detection of critical conditions such as arrhythmias and hypoxia. However, conventional cloud-centric architectures introduce significant end-to-end latency, which can compromise timely clinical response in safety-critical scenarios. Mobile Edge Computing (MEC) mitigates this limitation by bringing computation closer to data sources; yet, existing scheduling approaches remain largely system-centric and do not adequately incorporate patient-specific clinical risk into their decision-making. To address this gap, this paper presents CRAI-LCS, a clinical risk-aware Reinforcement learning (RL) framework for latency-constrained scheduling in HIoT systems. Unlike prior simulation-driven studies, CRAI-LCS integrates real physiological data from the PhysioNet MIT-BIH Arrhythmia database to construct realistic, data-driven workloads. Specifically, electrocardiogram (ECG) signals are segmented into time-windowed tasks with clinically grounded characteristics, including input size, computational demand, and urgency-aware deadlines. The framework combines data-driven clinical risk estimation, deadline-violation prediction, and RL-based scheduling to dynamically prioritize high-risk tasks while efficiently managing system resources. Experimental results demonstrate that CRAI-LCS consistently outperforms baseline approaches in terms of latency, deadline compliance, and resource utilization under realistic workload conditions. Ablation studies further confirm the individual contributions of the clinical risk-awareness and predictive scheduling components. Overall, these findings highlight the importance of incorporating real physiological data into scheduling design, providing a more reliable and clinically relevant foundation for next-generation healthcare edge intelligence systems.

J. B. Bin Jumah, H. Mirghani, Saad Alateeq et al. · 0 citations
Conference Jul 2026

DAStream: Efficient Drift-Adaptive Anomaly Detection for Streaming Data in Resource-Constrained Edge Intelligence

Real-time anomaly detection in industrial IoT (IIoT) often requires processing continuous data streams on resource-constrained edge nodes while addressing non-stationary data distributions caused by changes in device operating states. Traditional statistical or distance-based methods typically rely on fixed thresholds or static models, making it hard to maintain stable performance under complex conditions. Deep learning approaches are computationally intensive, making them unsuitable for resource-constrained edge devices. Existing methods struggle to balance dynamic adaptability to data distributions with computational efficiency. This paper proposes DAStream, a streaming anomaly detection method for edge environments. We propose a dual-phase online clustering framework with statistical enhancement to improve model stability. We design an adaptive anomaly scoring method that uses time-varying statistics to capture dynamic drifts in the data distribution, combined with an online-updated normalized deviation metric. We also propose a lightweight self-calibrating discrimination mechanism to enable dynamic decision boundaries. DAStream relies solely on recursive statistical computations, ensuring constant computational complexity and enabling efficient deployment on devices such as FPGAs. Experiments on typical IIoT datasets and an FPGA platform show that DAStream achieves a $\text{6 2. 6 \%}$ increase in throughput and a 40.7% reduction in energy consumption while maintaining detection performance comparable to the state-of-the-art method, validating its effectiveness and feasibility for IIoT edge intelligence scenarios.

Xiao Liu, Shubo Liu, Zhaohui Cai et al. · 0 citations
Open access Aug 2026

Adaptive Latency-Aware Predictive Framework for Mobile Edge Systems Using Dynamic Feature Reduction and Intelligent Scheduling

Low-latency and resource-efficient predictive analytics are essential for mobile edge computing applications, including smartphone-based human activity recognition (HAR). This study presents the Adaptive Latency Prediction and Scheduling (ALPS) framework, which aims to minimize inference latency, hardware energy consumption, and computational overhead while maintaining predictive accuracy. The ALPS framework comprises three primary modules: an adaptive principal component analysis (PCA) mechanism for real-time feature reduction, a lightweight ridge classifier optimized with an L regularization loss 2 function for rapid multi-class activity prediction, and a dynamic task scheduler that minimizes a combined latency-energy cost function to allocate processing tasks across mobile, edge, and cloud layers. Evaluation on the high-dimensional UCI-HAR smartphone dataset (10,299 samples, 561 features) demonstrated that the adaptive feature reduction module reduced the feature space to 68 principal components, resulting in an 87.9% reduction in dimensionality while maintaining 95% cumulative explained variance. Relative to conventional standalone models and isolated optimization baselines, ALPS achieved a 32% reduction in average end-to-end latency (120 ms), a 25% decrease in energy consumption (180 J), and a classification accuracy of 96.4% across six physical activities. The primary contribution of this study is the unified integration of adaptive data compression and distributed infrastructure scheduling into a scalable and energy-efficient pipeline for real-time edge intelligence.

M. Nohong, Nora’asikin Abu Bakar, Siti Roshaida Abd Razak et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.