Skip to content
#software testing Dataset Open access

Data for study "The Storage Condition at Measurement Time Decides Denoiser Selection for Frozen Deep Pedestrian Detectors"

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Basic information---------------------------1. Journal article: The Storage Condition at Measurement Time Decides Denoiser Selection for Frozen Deep Pedestrian Detectors 2. DOI: concept DOI: 10.5281/zenodo.22643935 (resolves to the latest version) version 1: 10.5281/zenodo.22643936 (2026-09-07, files restricted) version 2 DOI: 3. Contact information Name: Vo Thanh Kiet Institution: VSB - Technical University of Ostrava E-mail: kiet.vo.thanh.st@vsb.cz ORCID: https://orcid.org/0009-0002-3278-8755 4. Dataset publication date: version 1: 2026-09-07; version 2: 5. Place of publication: Ostrava, Czechia ------------------------------------------------------------------6. Dataset Description================================================================================ This dataset contains the measurements, derived tables, statistical test outputs, dated analysisprotocols, frame hashes, own-trained denoiser checkpoints, the frozen detector of the core grid andthe source code behind the article named in item 1. The study asks under which storage condition afull-reference fidelity metric has to be measured before its choice of denoiser can be trusted, whenthe denoiser feeds a pedestrian detector whose weights are frozen and the frames carry additiveGaussian noise followed by a lossy Joint Photographic Experts Group (JPEG) storage stage. Its answeris that restoration fidelity, measured on uncompressed images, reads as a property of the denoiser,but in front of a frozen detector it is also a property of the storage condition under which it ismeasured.This version also holds, for the INRIA, CrowdHuman and CityPersons arms, the per-image detectorpredictions of a re-inference of their 367 detector cells and the paired frame-bootstrap confidenceintervals computed from them (gpu_runs/g1_perimage/). Selection regret is the detection cost, in mean average precision at intersection-over-union (IoU)0.50 (mAP@50; person AP@50 on INRIA, CrowdHuman and CityPersons), of deploying the denoiser a metricranks first instead of the best candidate of the same roster at the same operating point. Measuredon the frames the detector receives, the peak signal-to-noise ratio (PSNR) picks a near-optimaldenoiser; measured benchmark-style on the uncompressed pair it costs far more; on the losslesscontrols the two coincide by definition, so the cost comes from the mismatch between measurement anddeployment, not from the metric. An offline noise-level (sigma-hat) routing table and Restormer'spixel pass-through show the same dependence (values under "Numbers of the abstract"). PSNR needs aclean reference, so the selection is an offline step on frames in the deployment's storagecondition. All degradations are injected; natively degraded footage is not covered. Arms. An arm is one dataset under one storage recipe, evaluated as a unit over its noise levels.Eight are lossy -- the PnPLO (People and Person-Like Objects) main grid and PnPLO arm 1 at JPEGquality 95, the quality-sweep arms at 85 and 75 (E3), the PnPLO sigma = 50 arm (E4), and INRIA,CityPersons and CrowdHuman at quality 95 -- and two are lossless controls, PnPLO arm 1 and INRIA,whose noisy frames are stored as Portable Network Graphics (PNG). The benchmark-style measurementexists on both INRIA arms, both PnPLO arm-1 arms and the two E3 arms; CityPersons and CrowdHuman areJPEG only. Design. The core grid uses the 235-image PnPLO test split (944 / 160 / 235 train/val/test) and afrozen YOLOv9-M detector (mAP@50 0.745 clean; 0.695 / 0.542 / 0.399 noisy at sigma = 10 / 20 / 30).Stage order: Gaussian noise, JPEG quality-95 re-encode, the denoiser (output stored again as JPEGquality 95 on the main grid, E4, CrowdHuman and CityPersons), the frozen detector. The lever is thatJPEG stage, switched with all weights untouched (PnPLO arm 1, INRIA); the switch changes the framethe detector receives and does not isolate the storage stage as the sole cause. Twelve restorers:sigma-conditioned SwinIR, SCUNet and FFDNet; the level-blind Restormer (Gaussian-trained) andPromptIR (all-in-one); NAFNet (real noise, the Smartphone Image Denoising Dataset, SIDD); SCUNet-real; the own-trained DnCNN, convolutionalautoencoder (CAE) and CAE-PSO (particle swarm optimization); BM3D; a Gaussian filter. SwinIR andSCUNet run at the nearest available level (15 / 25 / 25 at sigma = 10 / 20 / 30, 50 at sigma = 50);FFDNet, the own-trained and the classical methods take the true sigma. Metrics: d_mAP50 (Delta-d inthe analysis files; the within-sigma mAP@50 change against the noisy baseline of the same arm),PSNR, the structural similarity index measure (SSIM) and the Structural Preservation Score(SPS_grad, SPS_edge), a supporting screen tested by the pre-specified gate V1 to V4; thestorage-condition claim rests on the selection regret, the noise-level table and the JPEG ablation.External arms: INRIA Person (nine detectors trained on clean INRIA, 90 test frames), CrowdHuman (500validation frames, four detectors pretrained on Common Objects in Context (COCO), zero-shot) andCityPersons (a fixed 200-frame subset, three COCO detectors, a pre-specified replication gate). Relation to version 1--------------------------------------------------------------------------------Version 1 (10.5281/zenodo.22643936) accompanied an earlier manuscript under an earlier title andclaim axis, which the article of this version withdraws: under lossy storage the level-blindrestorers are the LOWER-fidelity ones, and SPS is a supporting screen, not a contribution. Everyversion-1 data, result, checkpoint, smoke and code file is in this version at its version-1path, except that readme.txt and LICENSE.txt are rewritten, manifest/MANIFEST.csv and00_PATHS_README.txt regenerated, three figure scripts (make_p14_figs_v2.py, p14_stats_fig2.py,make_p14_f8.py) replaced by the copies that drew the current figures (wording and layout only; noinput, path or computation changed), and the four files of the version-1 protocol/ folder replacedby the redacted first-commit protocol set (the same three documents and six more). New:analysis_v2/, gpu_kits/ (the re-inference kit, called G1 in file names), gpu_runs/(its outputs), frames/, reports/, the protocol/ set,results/sigma_hat_imgset_identity.json, verify_jpeg_chroma_420.py with its output, and00_SANITISATION.csv. Sanitisation. Shipped code is as run, with local paths made relative or replaced by tokens such as (explained in 00_PATHS_README.txt), internal names replaced by role words ("theanalyst", "the first author"), authorship clauses reduced to their date and working notes removed,so it is not byte-identical to the files executed; a few data and provenance files were tokenised the same way.00_SANITISATION.csv gives the Secure Hash Algorithm 256-bit (SHA-256) hash of the original and ofthe shipped copy of every processed file; a hash recorded inside the package is that of theORIGINAL. smoke/task_b_permutation_LOCAL.json had its machine label neutralised and its line endingsset to Unix line feeds (LF) (version-1 SHA-256 793b8db20a0d4a2e368f9ed0ac16e1d852cfaae42671e6d8e7cfe25141a216c2; norow in 00_SANITISATION.csv). checkpoints/detector_yolov9m_clean/weights/best.pt is byte-identical toversion 1; its pickled training arguments keep the authors' Colab training paths, which thesanitised args.yaml holds as tokens. Not included-------------------------------------------------------------------------------- * Image data: the four corpora (PnPLO, INRIA Person, CrowdHuman, CityPersons), the noisy and the restored sets; frames/ names every input frame with its SHA-256. * The ground-truth boxes of INRIA, CrowdHuman and CityPersons (see "Parquet release"). * Third-party restorer checkpoints and code (SwinIR, SCUNet, SCUNet-real, FFDNet, Restormer, PromptIR, NAFNet); the further PnPLO and the nine INRIA detectors (available from the authors upon reasonable request); ultralytics downloads the COCO detectors. * Graphics processing unit (GPU) kits G2 to G4 and two unfinished central processing unit (CPU) analyses (nothing in the article rests on them), material for an unused figure, the parent-study notebook (core-grid noise, own-trained training), and the smoke outputs and caches of the version-2 analyses. What this version does NOT hold-------------------------------------------------------------------------------- * Per-image detector predictions for any PnPLO arm (main grid, arm 1, E3, E4, arms 2, 3 and 3b, cross-detector sweep): only set-level detection values; the G1 kit does not re-infer them. * Per-image fidelity and structure of the nine corrected own-trained cells (set-level only); their rows in data/pnplo/p14_sps_per_image.csv are the legacy tiling pass. * Per-image SSIM and SPS on the external arms (per-image PSNR of Restormer, PromptIR and SwinIR is in analysis_v2/p14_passthrough_audit_per_image.csv), and per-image latency timings (mean, median, standard deviation and count only). Per-image re-inference of the external arms (gpu_runs/g1_perimage/)--------------------------------------------------------------------------------After the analyses of the article were fixed, every detector cell of the stored external arms wasre-run with the kit on the stored frames, along each arm's original evaluation path and ultralyticspin, keeping every prediction: INRIA lossless (inria_png) and JPEG (inria_jpeg), 9 detectors x 13conditions each; CityPersons 3 x 19; CrowdHuman 4 x 19 (clean rows included); 367 cells. Nodenoising was run. env.json records five sessions on 2026-09-26, one NVIDIA A100 (40 GB) each,with software versions, GPU and script hashes: a kit-1.4 session (stopped by its tripwire at itsseventh cell; six cells kept), three kit-1.7 sessions and one kit-1.8 diagnosis of a single cell.The kit version of each cell is in the p14_g1_provenance record of its parquet file; the diagnosisse

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#computer vision Review Mar 2008

Agile methods in European embedded software development organisations: a survey on the actual use and usefulness of Extreme Programming and Scrum

The results show that the embedded industry has been able to apply agile methods in its development processes and that the appreciation of the agile methods and their individual practices appears to increase once adopted and applied in practice.

O. Salo, P. Abrahamsson · 238 citations · ⚡9
#computer vision Open access Jul 2017

What happens when software developers are (un)happy

Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 236 citations · ⚡13
#computer vision Open access Oct 2004

Mobile-D: an agile approach for mobile application development

The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.

P. Abrahamsson, Antti Hanhineva, H. Hulkko et al. · 225 citations · ⚡18

Related blog posts

MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.