Synthetic Dataset for PRIVAGEN-SHIELD: Privacy-Preserving Genomic Anomaly Detection
Abstract
Dataset Description This repository contains the synthetic genomic workflow dataset generated and used for the experimental evaluation of the PRIVAGEN-SHIELD framework, “PRIVAGEN-SHIELD: Privacy-Preserving Genomic Anomaly Detection via Hierarchical Blockchain-Authenticated Merkle Trees and Local Differential Privacy.” The dataset is entirely synthetic and does not contain data from human participants, patients, or identifiable individuals. The complete dataset contains 50,000 synthetic genomic workflow records generated using a fixed random seed of 42, with an anomaly rate of 5%. Each record contains synthetic CRISPR workflow information, including edit identifier, CRISPR log, phenotype outcome, blockchain hash, timestamp, gene, genomic position, editing type, DNA sequence, and anomaly category. The synthetic DNA sequences are randomly generated using the nucleotide bases A, C, G, and T. Eight anomaly categories are represented: malicious edit, extreme low phenotype, temporal forward anomaly, hash collision, high-risk combination, temporal backward anomaly, unauthorized gene, and combined anomaly. For experimental evaluation, the chronologically ordered dataset was divided into 60% training, 20% validation, and 20% testing subsets. The accompanying documentation provides information required to understand and reproduce the dataset-generation and experimental workflow. Data type: Synthetic genomic/workflow dataNumber of records: 50,000Anomaly rate: 5%Random seed: 42Dataset generation: Python-based synthetic data generator This dataset is intended for research, benchmarking, reproducibility, and evaluation of privacy-preserving genomic anomaly detection methods.