Skip to content
Book

Decay Driven Multi Objective Optimization for HPC nodes

Jul 2026 · IEEE International Symposium on High-Performance Parallel Distributed Computing · pp. 584-585 · 0 citations · 11 references
Computer Science

TL;DR

A decay-driven preference formulation is presented that replaces abstract preference weights with a user-defined performance degradation parameter, and leverages a domain-specific monotonic relationship between application progress and power cap, removing the need for interpolation models typically required in preference-driven MORL.

Abstract

Energy efficiency has become a primary design constraint in modern high-performance computing systems, where energy costs increasingly dominate total cost of ownership. At the same time, applications are subject to strict performance requirements, making energy optimization inherently a multi-objective problem involving conflicting objectives of minimizing energy consumption while preserving application performance [1, 2]. Hardware interfaces such as Intel RAPL enable runtime power capping, allowing dynamic control over processor power budgets, but selecting optimal operating points remains challenging due to application-dependent and non-linear performance behavior [5, 7]. This work presents Decay-Driven Multi-Objective Reinforcement Learning (DDMORL), an offline preference-driven reinforcement learning framework for adaptive power control in HPC systems. The proposed approach extends preference-driven multi-objective reinforcement learning [3] to a fully offline setting [9], eliminating the need for unsafe online exploration and enabling deployment in production environments. Unlike prior approaches that rely on fixed scalarization or multiple trained policies, DDMORL learns a single preference-conditioned controller that spans the entire energy performance trade-off space. A key contribution of this work is a decay-driven preference formulation that replaces abstract preference weights with a user-defined performance degradation parameter. Users specify a maximum tolerable decay, which is then analytically mapped to a preference vector aligned with the corresponding operating point on the Pareto front. This mapping leverages a domain-specific monotonic relationship between application progress and power cap, removing the need for interpolation models typically required in preference-driven MORL.

View source

Similar papers

Open access Aug 2026

Edge-cloud collaboration-driven predictive modeling for high-performance computing centers.

This study proposes a deployment-oriented edge-cloud collaboration (ECC) framework integrated with a transient-aware predictive architecture, named FS-Attention, designed to balance transient responsiveness, engineering deployability, and decision transparency, which achieves competitive full-year prediction accuracy.

Shuai-Yin Ma, Ye-Ye Cao, Yang Liu et al. · 0 citations
Book Open access Jul 2026

Extending the Life of HPC Systems in Resource Constrained Environments: Mapping Productivity-Energy Trade-offs in Memory-Bound Workloads via DVFS and Core Scaling on Repurposed Hardware

This work examines the impact of limiting active cores on repurposed nodes and introduces deep C-state power-gating, fully saturated workloads, and hardware-level power measurements to address viability in complex applications such as OpenFOAM.

Bryan Johnston, Suné Toerien, Vele Nefale et al. · 0 citations
Book Open access Jul 2026

Toward an Integrated Theory of Adaptive Scheduling in High Performance Computing: A Queuing-Theoretic and Computational Learning Perspective.

High Performance Computing (HPC) systems increasingly operate under heterogeneous workloads, dynamic resource availability, and stringent performance and energy constraints. Traditional batch scheduling policies such as First-Come First-Served (FCFS), backfilling, and priority-based heuristics rely on static assumption...

Rodgers Kimera, Ali Najib, David Kakeeto · 0 citations
Open access 2026

Energy-Aware Task Scheduling using Grasshopper Optimisation Algorithm

Efficient energy scheduling in heterogeneous computing environments is a critical challenge, as task allocation decisions directly affect both energy consumption and execution performance. This work presents an energy aware scheduling framework based on a discretized grasshopper optimization algorithm (GOA), designed t...

Macauley Opuwari, C. Igiri, D. Ikeh · 0 citations
Open access Jul 2026

Comparative Performance and Computational Complexity Analysis of Hybrid WGO–DRL and Heuristic, Metaheuristic and Bio-Inspired Scheduling Algorithms in Cloud Computing

A hybrid scheduling framework that integrates Hybrid Wild Goose Optimization (HWGO) with Deep Reinforcement Learning (DRL) is investigated, indicating that intelligent hybrid optimization techniques can provide adaptive and efficient task scheduling solutions for modern cloud computing environments.

Annaiah H, A. Rajesh · 0 citations
Open access Aug 2026

EMC+: An Opportunistic Elasticity Method for Improving System Throughput and CPU Utilization in Cloud Data Centers

The new EMC+ proposal is an OS‐driven elasticity manager for container‐based environments that continuously estimates idle core cycles left by regular (inelastic) applications, and reallocates idle cores to elastic ones, even during short time intervals, and has minimal impact on the performance and QoS of colocated in...

J. C. Saez, Carlos Bilbao, Manuel Prieto-Matías · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.