Skip to content
#graph neural networks Dataset Open access

IMOS-SWIL District Heating Dataset

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)
Integrated Energy Systems Optimization

Abstract

This dataset provides synchronized, high-resolution multi-physics measurements collected by IMOS-EPFL in collaboration with Aalborg University using the controlled district heating testbed at the Smart Water Infrastructures Laboratory (SWIL), Aalborg University, Denmark. The data were collected as part of the Intelligent Thermal Energy Networks (ITEN) research project, funded by the Swiss Federal Institute of Metrology (METAS), and were developed to support research on smart heat meter monitoring, data-driven monitoring, and machine learning for district heating systems. The experimental system consists of a heat station, pumping station, pipe units, and two consumer substations arranged in a representative district heating configuration. Each consumer station is equipped with two redundant commercial smart heat meters installed in series. Their measurements are averaged to reduce measurement noise and provide reliable reference signals for virtual sensing evaluation. The release contains four CSV files corresponding to four distinct operating conditions: File Time samples Approx. duration DS01.csv 15,218 8.45 h DS02.csv 15,286 8.49 h DS03.csv 15,078 8.38 h DS04.csv 8,200 4.56 h Total 53,782 29.88 h All variables are synchronized and sampled at 0.5 Hz, corresponding to one measurement every two seconds. Each CSV contains 65 synchronized variables and contains no missing values. SCADA measurements The dataset includes measurements from four main parts of the testbed: Heat station: variables with the prefix Boiler_ Pipe unit: variables with the prefix Pipe_ Pumping station: variables with the prefix Pump_ Consumer units: variables with the prefix Consumer_ The principal SCADA variables used in the accompanying study consist of 44 sensor signals covering three physical modalities: 20 temperature measurements, distributed across the heat station, pipe units, pumping station, and consumer units; 15 pressure measurements, including pressure and differential-pressure measurements across the network; 9 flow-rate measurements, located at the heat station, pipe units, pumping station, and consumer units. The complete list of SCADA inputs and their locations is reported in the associated publication. Typical variable names include Boiler_T1, Pipe_T2_1, Consumer_T_21, and Pump_T3_1 for temperature measurements; Boiler_dP2, Pipe_P1_1, Consumer_P_21, and Pump_P3_1 for pressure measurements; and Boiler_Q2_1, Pipe_Q3_1, and Consumer_Q1_21 for flow measurements. The released CSV files additionally retain auxiliary experimental/testbed channels beyond the 44 SCADA variables used as model inputs in the accompanying study. Smart-meter reference measurements In addition to the SCADA measurements, each CSV contains 10 processed smart-meter variables obtained by averaging measurements from the redundant meter pairs at the two consumer substations. For Consumer / Smart Meter location 1, the following variables are provided: avg_of_F1_F2 — averaged flow measurement; avg_of_V1_V2 — averaged volume measurement; avg_of_E1_E2 — averaged energy measurement; avg_of_T1_in_T2_in — averaged inlet temperature; avg_of_T1_out_T2_out — averaged outlet temperature. For Consumer / Smart Meter location 2, the corresponding variables are: avg_of_F3_F4 — averaged flow measurement; avg_of_V3_V4 — averaged volume measurement; avg_of_E3_E4 — averaged energy measurement; avg_of_T3_in_T4_in — averaged inlet temperature; avg_of_T3_out_T4_out — averaged outlet temperature. The commercial smart meters provide high-accuracy consumer-side measurements and are used as reference measurements in the accompanying study. The virtual sensing experiments reported in the associated paper focus on six reference variables: flow rate, inlet temperature, and outlet temperature at each of the two consumer substations. The volume and energy channels are retained in this public release to enable additional research applications. Operating conditions The four datasets correspond to different consumer-demand configurations. Consumer fan speeds were varied between experiments, producing four combinations of operating conditions. In addition, boiler on-off control introduced thermal variability by reheating the supply water when its temperature decreased and switching off the boiler after the target temperature was restored. These interventions generate different hydraulic and thermal operating regimes. The four datasets are therefore suitable for evaluating generalization across operating regimes. In the accompanying paper, a leave-one-dataset-out evaluation is used: three operating conditions are used for training, and the remaining condition is used for testing. Intended use The dataset can support research on: virtual sensing and soft sensing; virtual smart heat metering; district-heating monitoring and state estimation; graph neural networks and spatial-temporal learning; multivariate and multi-physics time-series modeling; sensor reconstruction and missing-measurement estimation; anomaly and sensor-fault detection; data-driven digital twins; thermo-hydraulic system identification and monitoring. A key motivation for releasing this dataset is the limited availability of public district heating datasets that combine synchronized temperature, pressure, flow, and consumer-side smart-meter measurements. The dataset is intended as a controlled and reproducible benchmark for developing and comparing data-driven virtual-sensing methods. Associated publication This dataset accompanies: K. Faghih Niresi, C. Møller Jensen, C. S. Kallesøe, R. Wisniewski, and O. Fink, “Virtual Smart Metering in District Heating Networks via Heterogeneous Spatial-Temporal Graph Neural Networks,” Energy and Buildings, 2026. Users of this dataset are kindly requested to cite both the dataset and the associated publication.

View source

Similar papers

#computer vision Conference Aug 2008

Scrum in a Multiproject Environment: An Ethnographically-Inspired Case Study on the Adoption Challenges

Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoptio...

A. Marchenko, P. Abrahamsson · 59 citations · ⚡11
#computer vision Open access Sep 2012

Making the leap to a software platform strategy: Issues and challenges

A comprehensive taxonomy of the challenges faced when a medium-scale organization decided to adopt software platforms is provided, namely: business challenges, organizational challenges, technical challenges, and people challenges.

Yaser Ghanam, F. Maurer, P. Abrahamsson · 41 citations · ⚡3
#machine learning Open access Mar 2024

Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction

MCGLPPI, a novel geometric representation learning framework that combines graph neural networks (GNNs) with the MARTINI molecular coarse-grained (CG) model to predict overall PPI properties accurately and efficiently, offers an effective and efficient solution for PPI overall property predictions.

Yang Yue, Shu Li, Yihua Cheng et al. · 15 citations

PepPCBench is a Comprehensive Benchmarking Framework for Protein-Peptide Complex Structure Prediction

PepPCBench enables a robust evaluation of PFNN-based methods and supports their continued development for peptide-protein structure prediction, and highlights the influence of peptide length, conformational flexibility, and training set similarity on prediction accuracy.

Si-Long Zhai, Huifeng Zhao, Ji-Ke Wang et al. · 13 citations · ⚡1
#machine learning Open access Sep 2025

Unified and explainable molecular representation learning for imperfectly annotated data from the hypergraph view

OmniMol is presented, a framework using hypergraphs to improve predictions of molecular properties, addressing challenges of imperfect data annotation and enhancing model explainability, and achieves state-of-the-art performance in properties prediction.

Bowen Wang, Junyou Li, Donghao Zhou et al. · 11 citations

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.