Data for study "A Structural Preservation Framework for Denoiser Selection in YOLO-Based Pedestrian Detection under Sensor Noise"
Abstract
Basic information---------------------------1. Journal article: The Storage Condition at Measurement Time Decides Denoiser Selection for Frozen Deep Pedestrian Detectors 2. DOI: concept DOI: 10.5281/zenodo.22643935 (resolves to the latest version) version 1: 10.5281/zenodo.22643936 (2026-09-07, files restricted) version 2 DOI: 10.5281/zenodo.23044857 3. Contact information Name: Vo Thanh Kiet Institution: VSB - Technical University of Ostrava E-mail: kiet.vo.thanh.st@vsb.cz ORCID: https://orcid.org/0009-0002-3278-8755 4. Dataset publication date: version 1: 2026-09-07; version 2: 5. Place of publication: Ostrava, Czechia ------------------------------------------------------------------6. Dataset Description================================================================================ This dataset contains the measurements, derived tables, statistical test outputs, dated analysisprotocols, frame hashes, own-trained denoiser checkpoints, the frozen detector of the core grid andthe source code behind the article named in item 1. The study asks under which storage condition afull-reference fidelity metric has to be measured before its choice of denoiser can be trusted, whenthe denoiser feeds a pedestrian detector whose weights are frozen and the frames carry additiveGaussian noise followed by a lossy Joint Photographic Experts Group (JPEG) storage stage. Its answeris that restoration fidelity, measured on uncompressed images, reads as a property of the denoiser,but in front of a frozen detector it is also a property of the storage condition under which it ismeasured.This version also holds, for the INRIA, CrowdHuman and CityPersons arms, the per-image detectorpredictions of a re-inference of their 367 detector cells and the paired frame-bootstrap confidenceintervals computed from them (gpu_runs/g1_perimage/). Selection regret is the detection cost, in mean average precision at intersection-over-union (IoU)0.50 (mAP@50; person AP@50 on INRIA, CrowdHuman and CityPersons), of deploying the denoiser a metricranks first instead of the best candidate of the same roster at the same operating point. Measuredon the frames the detector receives, the peak signal-to-noise ratio (PSNR) picks a near-optimaldenoiser; measured benchmark-style on the uncompressed pair it costs far more; on the losslesscontrols the two coincide by definition, so the cost comes from the mismatch between measurement anddeployment, not from the metric. An offline noise-level (sigma-hat) routing table and Restormer'spixel pass-through show the same dependence (values under "Numbers of the abstract"). PSNR needs aclean reference, so the selection is an offline step on frames in the deployment's storagecondition. All degradations are injected; natively degraded footage is not covered. Arms. An arm is one dataset under one storage recipe, evaluated as a unit over its noise levels.Eight are lossy -- the PnPLO (People and Person-Like Objects) main grid and PnPLO arm 1 at JPEGquality 95, the quality-sweep arms at 85 and 75 (E3), the PnPLO sigma = 50 arm (E4), and INRIA,CityPersons and CrowdHuman at quality 95 -- and two are lossless controls, PnPLO arm 1 and INRIA,whose noisy frames are stored as Portable Network Graphics (PNG). The benchmark-style measurementexists on both INRIA arms, both PnPLO arm-1 arms and the two E3 arms; CityPersons and CrowdHuman areJPEG only. Design. The core grid uses the 235-image PnPLO test split (944 / 160 / 235 train/val/test) and afrozen YOLOv9-M detector (mAP@50 0.745 clean; 0.695 / 0.542 / 0.399 noisy at sigma = 10 / 20 / 30).Stage order: Gaussian noise, JPEG quality-95 re-encode, the denoiser (output stored again as JPEGquality 95 on the main grid, E4, CrowdHuman and CityPersons), the frozen detector. The lever is thatJPEG stage, switched with all weights untouched (PnPLO arm 1, INRIA); the switch changes the framethe detector receives and does not isolate the storage stage as the sole cause. Twelve restorers:sigma-conditioned SwinIR, SCUNet and FFDNet; the level-blind Restormer (Gaussian-trained) andPromptIR (all-in-one); NAFNet (real noise, the Smartphone Image Denoising Dataset, SIDD); SCUNet-real; the own-trained DnCNN, convolutionalautoencoder (CAE) and CAE-PSO (particle swarm optimization); BM3D; a Gaussian filter. SwinIR andSCUNet run at the nearest available level (15 / 25 / 25 at sigma = 10 / 20 / 30, 50 at sigma = 50);FFDNet, the own-trained and the classical methods take the true sigma. Metrics: d_mAP50 (Delta-d inthe analysis files; the within-sigma mAP@50 change against the noisy baseline of the same arm),PSNR, the structural similarity index measure (SSIM) and the Structural Preservation Score(SPS_grad, SPS_edge), a supporting screen tested by the pre-specified gate V1 to V4; thestorage-condition claim rests on the selection regret, the noise-level table and the JPEG ablation.External arms: INRIA Person (nine detectors trained on clean INRIA, 90 test frames), CrowdHuman (500validation frames, four detectors pretrained on Common Objects in Context (COCO), zero-shot) andCityPersons (a fixed 200-frame subset, three COCO detectors, a pre-specified replication gate). Relation to version 1--------------------------------------------------------------------------------Version 1 (10.5281/zenodo.22643936) accompanied an earlier manuscript under an earlier title andclaim axis, which the article of this version withdraws: under lossy storage the level-blindrestorers are the LOWER-fidelity ones, and SPS is a supporting screen, not a contribution. Everyversion-1 data, result, checkpoint, smoke and code file is in this version at its version-1path, except that readme.txt and LICENSE.txt are rewritten, manifest/MANIFEST.csv and00_PATHS_README.txt regenerated, three figure scripts (make_p14_figs_v2.py, p14_stats_fig2.py,make_p14_f8.py) replaced by the copies that drew the current figures (wording and layout only; noinput, path or computation changed), and the four files of the version-1 protocol/ folder replacedby the redacted first-commit protocol set (the same three documents and six more). New:analysis_v2/, gpu_kits/ (the re-inference kit, called G1 in file names), gpu_runs/(its outputs), frames/, reports/, the protocol/ set,results/sigma_hat_imgset_identity.json, verify_jpeg_chroma_420.py with its output, and00_SANITISATION.csv. Sanitisation. Shipped code is as run, with local paths made relative or replaced by tokens such as (explained in 00_PATHS_README.txt), internal names replaced by role words ("theanalyst", "the first author"), authorship clauses reduced to their date and working notes removed,so it is not byte-identical to the files executed; a few data and provenance files were tokenised the same way.00_SANITISATION.csv gives the Secure Hash Algorithm 256-bit (SHA-256) hash of the original and ofthe shipped copy of every processed file; a hash recorded inside the package is that of theORIGINAL. smoke/task_b_permutation_LOCAL.json had its machine label neutralised and its line endingsset to Unix line feeds (LF) (version-1 SHA-256 793b8db20a0d4a2e368f9ed0ac16e1d852cfaae42671e6d8e7cfe25141a216c2; norow in 00_SANITISATION.csv). checkpoints/detector_yolov9m_clean/weights/best.pt is byte-identical toversion 1; its pickled training arguments keep the authors' Colab training paths, which thesanitised args.yaml holds as tokens. Not included-------------------------------------------------------------------------------- * Image data: the four corpora (PnPLO, INRIA Person, CrowdHuman, CityPersons), the noisy and the restored sets; frames/ names every input frame with its SHA-256. * The ground-truth boxes of INRIA, CrowdHuman and CityPersons (see "Parquet release"). * Third-party restorer checkpoints and code (SwinIR, SCUNet, SCUNet-real, FFDNet, Restormer, PromptIR, NAFNet); the further PnPLO and the nine INRIA detectors (available from the authors upon reasonable request); ultralytics downloads the COCO detectors. * Graphics processing unit (GPU) kits G2 to G4 and two unfinished central processing unit (CPU) analyses (nothing in the article rests on them), material for an unused figure, the parent-study notebook (core-grid noise, own-trained training), and the smoke outputs and caches of the version-2 analyses. What this version does NOT hold-------------------------------------------------------------------------------- * Per-image detector predictions for any PnPLO arm (main grid, arm 1, E3, E4, arms 2, 3 and 3b, cross-detector sweep): only set-level detection values; the G1 kit does not re-infer them. * Per-image fidelity and structure of the nine corrected own-trained cells (set-level only); their rows in data/pnplo/p14_sps_per_image.csv are the legacy tiling pass. * Per-image SSIM and SPS on the external arms (per-image PSNR of Restormer, PromptIR and SwinIR is in analysis_v2/p14_passthrough_audit_per_image.csv), and per-image latency timings (mean, median, standard deviation and count only). Per-image re-inference of the external arms (gpu_runs/g1_perimage/)--------------------------------------------------------------------------------After the analyses of the article were fixed, every detector cell of the stored external arms wasre-run with the kit on the stored frames, along each arm's original evaluation path and ultralyticspin, keeping every prediction: INRIA lossless (inria_png) and JPEG (inria_jpeg), 9 detectors x 13conditions each; CityPersons 3 x 19; CrowdHuman 4 x 19 (clean rows included); 367 cells. Nodenoising was run. env.json records five sessions on 2026-09-26, one NVIDIA A100 (40 GB) each,with software versions, GPU and script hashes: a kit-1.4 session (stopped by its tripwire at itsseventh cell; six cells kept), three kit-1.7 sessions and one kit-1.8 diagnosis of a single cell.The kit version of each cell is in the p14_g1_provenance record of its parquet file; the diagnosissession's chec