Skip to content
#explainable ai Dataset Open access

Replication Data and Code for: DE-LU Diffusion Price Forecast

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Guangzhou College of Technology and BusinessSchool of Business · Guangzhou, Guangdong, China Replication data and code for the manuscript: In what context does generative AI create economic value? Diffusion forecasts, negative prices, and the exercise-breadth law of forecast value Submitted to Global Finance Journal, Special Issue on AI and Capital Markets. This record contains the complete replication package for the above manuscript: all code, all derived data artefacts and the authoritative aggregated result files. It holds 213 files (64 MB) in four entries at the archive root — code/, data/, _build/ and README.md — and is entirely in English. Every number, table and figure reported in the paper is derived from the files shipped here; none is transcribed by hand. What the study does The paper estimates a conditional denoising diffusion model of the German–Luxembourgish (DE-LU) day-ahead electricity price in the difference domain and races it against eight forecasting families under a common 19-quantile protocol. The sample is 122,543 hourly observations from 9 January 2012 to 31 December 2025; the test period is 2024–2025, covering 17,520 held-out hours (730 forecasting days). On average accuracy the model loses, both to its own level-domain twin and to quantile LightGBM. On the revenue that a 1 MW / 2 MWh storage position extracts under a tail-contingent threshold rule the ranking reverses. The paper then traces the boundary of that reversal with a one-parameter family of dispatch rules, reporting where the premium is invariant, where it decays monotonically and where it changes sign. Austria serves as a cross-market transfer check. The full 41-configuration experiment set runs on a single laptop GPU. How the package is organised code/ — all 58 Python scripts (85 files): the training layer, the analysis layer, the revision analyses and the deliverable assembly. data/ — the derived data artefacts (105 files): per-configuration metrics, prediction ensembles, trained checkpoints, data documentation and GPU records. _build/ — the authoritative aggregated result files (22 files): the single source of truth for every number in the manuscript. The analysis layer writes this directory; the manuscript, the figure and table manifest and the four web pages all read it through one loader, so definitional drift between them is impossible. README.md — top-level guide to the package, its contents, its reproducibility limits, its data sources and its licence. code/README.md is the authoritative reproduction entry point: environment preparation, execution order, a reproduction checklist of expected constants and the known limitations. Contents in detail code/train/ — 13 scripts. Model definitions, the training loop and the 41-configuration runner. exp_models.py defines the 1D U-Net conditional diffusion, the LSTM / GRU / Seq2Seq / Transformer baselines and LightGBM / LEAR. exp_common.py is the shared library: data loading, window splitting, metric computation and DDIM sampling. run_suite.py drives the 41 configurations, run_anchor.py the seven-member anchor family, run_seeds.py the three-seed by two-domain controlled comparison, and train_delu.py single-model training and evaluation. Six diag_*.py scripts produce the splice, anchor-strength, calibration and sampling diagnostics, and gpu_bench.py re-runs the VRAM and throughput sweep. code/analysis/ — 34 scripts. The evaluation layers, the figure layer and the assembly of every table. verify_metrics.py and build_summary.py aggregate the 41 configurations; analysis_layers.py produces the event-probability, calibration, economic and descriptive-panel layers; frontier.py, dilution.py, mechanism_chain.py, scorecard.py and the econ_*.py family build the mechanisms and robustness results; make_figures_a.py and make_figures_b.py generate the manuscript figures at 300 dpi; make_paper_v4.py, make_refs_docx.py and make_manifest.py assemble the Word deliverables; and four make_html_*.py scripts build the four single-file web reports. The _patch/ subfolder holds four archived one-off scripts that are not part of the pipeline. Shared modules include pdata.py (the one data-loading layer), docxio.py (a .docx reader/writer that does not need python-docx and supports inline hyperlinks, which is how every DOI in the reference list becomes clickable), equations.py (the single definition of the 20 numbered equations), figstyle.py and figmap.py (the figure style and the figure-number-to-file mapping). code/revision/ — 7 scripts plus 3 result files. Everything the manuscript reports that is not produced by the original pipeline: e1e2e3_experiments.py (Diebold–Mariano tests with Newey–West lag 23 and a stationary bootstrap; the three dispatch-rule families on one common accounting ledger including the predict-then-optimise linear programme and its perfect-foresight upper bound; and PIT recalibration across all eighteen core configurations), efficiency_audit.py (the economic layer recomputed under both efficiency conventions), flexibility_law.py (the exercise-breadth law across twenty-five breadths), and redraw_fig09.py, redraw_fig10.py and redraw_roadmap.py. latex_to_omml.py is the LaTeX-to-OMML converter used to typeset the equations as native Word objects; it reproduces no reported number. data/results_ext/ — 90 files. The inputs to every analysis: 45 JSON files of per-configuration metrics (41 configurations plus diagnostics), 41 NPZ prediction ensembles (the full 19-quantile forecast for each configuration over the 17,520 test hours), 3 trained PyTorch checkpoints, and the run log. data/processed/ — 8 files. Yearly and validation summary tables (CSV) and three JSON files backing the descriptive tables. data/meta/ — 2 documents. DATA_README.md is the data dictionary: sources, bidding-zone definition, an inventory of the raw files, the processed datasets and the measured conclusions on year availability. LAGO_ALIGNMENT_REPORT.md is the two-source splice alignment report for the 2012–2025 extended series, including the time-zone determination, the cross-source calibration of load and wind-and-solar forecasts, continuity at the splice boundary and the marker-column semantics. data/gpu/ — 5 files. The hardware record and the GPU benchmark logs (throughput, VRAM sweep, sampling benchmark, end-to-end runtime estimate). _build/ — 22 JSON and TXT files, listed and explained inside code/README.md: the aggregation, event, calibration, economic, panel, frontier, dilution, mechanism, robustness, scorecard, reference and audit artefacts, together with the cross-reference audit record. Reproducibility Two levels are distinguished deliberately, so that a reader knows exactly what is verified and what is not. Level 1 — every reported number. Re-derivable from the shipped artefacts without retraining and without re-downloading any market data. The authoritative _build/ directory is part of this record, so no reported number depends on regeneration. In addition, fourteen of the sixteen analysis-layer scripts run from this package alone and thirteen regenerate their output file, from data/results_ext/. The figures are regenerated by make_figures_a.py, make_figures_b.py and make_fig_gpu.py. Level 2 — retraining from raw inputs. Requires the three large hourly series (DE_LU_hourly.csv, AT_hourly.csv, DE_LU_extended_2012_2025.csv, about 320 MB), which are not included because they are rebuilt from the public sources listed below; no credentials are required. Three consequences, stated plainly. First, code/train/ cannot retrain the 41 configurations here, so the trained checkpoints, the per-configuration metric files and all 41 prediction ensembles are shipped instead. Second, _build/descriptives.json and the descriptive panel inside analysis_panel.json and summary.json['panel'] cannot be rebuilt, because descriptives.py and analysis_layers.panel_layer() read the extended series; the authoritative versions are shipped. Third, _build/refs.json cannot be rebuilt, because extract_refs.py reads a reference document belonging to the manuscript workflow; the authoritative 68-entry file is shipped. Computational environment The reported results were produced on a single laptop GPU, recorded in data/gpu/hw.json and data/gpu/gpu_bench.json: HP OMEN Gaming Laptop 16-am0xxx; Intel Core i7-14650HX, 16 cores and 24 threads; 64 GB DDR5-5600; NVIDIA GeForce RTX 5070 Laptop GPU with 8,151 MiB of VRAM (Blackwell, sm_120, driver 616.64); Microsoft Windows 11 Pro, build 26200. Software: Python 3.13.14 and PyTorch 2.14.0+cu130 with CUDA 13.0; the CUDA 13.0 build is required because the sm_120 kernels are absent from the default PyPI wheels. The full 41-configuration experiment set runs in about 35.6 minutes. torch.compile is not used anywhere in the project, because Triton support on Windows is incomplete and compilation either fails or silently falls back to eager mode. Reproduction checklist The constants below let a reader confirm that a reproduction agrees: 41 experimental configurations (main table 18, ablation 9, anchor 7, transfer 5, robustness 2); 122,543 observed hours; 17,520 test hours across 730 forecasting days; 4,003 treated and 2,922 control training windows; 19 quantiles at τ = 0.05 to 0.95; and CRPS identically equal to twice Pinball, at a ratio of 2.0000 for all 41 configurations, because CRPS is computed by the standard (2/M)ΣL formula. Known limitations are reported as they stand rather than smoothed away: the interval coverage of the crisis-excluded robustness configuration falls far below nominal; the method has no advantage for spike events, where the best configuration is not the one proposed here; and the revenue calculation is an idealised convention that excludes transmission constraints, ancillary services and capacity markets. code/README.md lists these in full, together with the two references for which the publishers assign no DOI and the audit that established this. Data so

View source

Similar papers

#artificial intelligence Conference Open access Apr 2020

ECCOLA - a Method for Implementing Ethically Aligned AI Systems

The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.

Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson · 64 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Open access Mar 2024

LLM-based agents for automating the enhancement of user story quality: An early report

The use of large language models to automatically improve the user story quality in Austrian Post Group IT agile teams is explored, with a reference model for an Autonomous LLM-based Agent System developed and implemented at the company.

Zheying Zhang, M. Rayhan, Tomas Herda et al. · 48 citations · ⚡4
#computer vision Review Mar 2024

System for systematic literature review using multiple AI agents: Concept and an empirical evaluation

This paper introduces a novel multi-AI-agent system designed to fully automate SLRs, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision.

Abdul Malik Sami, Z. Rasheed, Kai-Kristian Kemell et al. · 44 citations · ⚡2
#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 41 citations
#artificial intelligence Conference Open access Jun 2018

The Key Concepts of Ethics of Artificial Intelligence

It is suggested that the focus on finding keywords is the first step in guiding and providing direction for future research in the AI ethics field.

Ville Vakkuri, P. Abrahamsson · 39 citations · ⚡2

Related blog posts

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.