Skip to content

HUNTINGTON'S DISEASE ENRICHED COHORT

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This Zenodo record contains a free, evaluation-grade sample of 10,000 synthetic patient records — a genuine, unmodified subset of the full production dataset, not a separate re-generation. It is provided so that researchers, ML engineers, and clinical data scientists can inspect the schema, phenotypic depth, and biomarker architecture before licensing the complete asset. The full commercial dataset — 5,000,000 synthetic patients, 11-year longitudinal depth (2016–2026), complete multi-ethnic population stratification, and the full pedigree network — is a licensed product and is not included in this record. Full dataset & enterprise licensing: https://sentineldata.com.ua/dataset/huntington-s-disease-cohort-dataset Abstract This dataset is a synthetic, privacy-safe electronic health record (EHR) cohort standardized to the OMOP Common Data Model (CDM) v5.4, built to address a persistent problem in rare-disease machine learning: the near-total absence of large, structurally complete, freely shareable datasets for orphan neurodegenerative conditions. The cohort is enriched around Huntington's Disease (HD), combining longitudinal clinical trajectories, genetic biomarker assertions (CAG repeat length), multi-generational pedigree graphs, and a realistic comorbidity background. No real patient data of any kind is contained in this file — every record is algorithmically generated. Clinical & Scientific Background Huntington's disease is an autosomal-dominant neurodegenerative disorder caused by an expanded CAG trinucleotide repeat in the HTT gene on chromosome 4p16.3, producing a toxic polyglutamine tract in the huntingtin protein and progressive striatal and cortical degeneration [1,7]. Clinically it presents as a triad of chorea, cognitive decline, and psychiatric disturbance, typically manifesting between 30 and 50 years of age, with juvenile-onset forms appearing in carriers of markedly longer repeats [6,7]. Repeat-length genetics follow defined clinical thresholds: alleles of 27–35 repeats are intermediate (unstable in transmission but not disease-causing in the carrier); 36–39 repeats are classified as reduced-penetrance; and 40 or more repeats are fully penetrant, virtually guaranteeing disease within a normal lifespan [1,2]. This threshold structure is the biological backbone against which any HD genetic dataset should be evaluated. Epidemiologically, HD is a genuinely rare disease. Pooled global prevalence is estimated at 3.9–4.9 per 100,000 persons, rising to 5.6–8.6 per 100,000 in populations of European descent (Europe, North America, Oceania) [3,4,5]. That scarcity is precisely why real-world HD datasets large enough for robust ML work are almost impossible to assemble without pooling many clinical sites — and why synthetic augmentation has scientific and commercial value. Why Synthetic Data — The Rare-Disease Gap Three structural problems make HD, and orphan diseases generally, a poor fit for conventional EHR data science: Prevalence is too low for statistical power. At roughly 5–8 cases per 100,000, a random hospital extract of even 1 million patients yields on the order of 50–80 HD cases — too few to train or validate most supervised models. Real HD registries are access-restricted. Data with genetic testing results and pedigree information is highly identifying; legitimate registries (e.g. Enroll-HD) require credentialed access and cannot be freely redistributed. Prodromal and family-structure data are the hardest signal to obtain. The features researchers most want — psychiatric prodrome, inheritance vectors, biomarker trajectories — are exactly the features most restricted in real data due to re-identification risk. Synthetic generation removes the access barrier entirely: no consent, no IRB, no re-identification risk, while preserving the statistical shape of the disease. Cohort Construction & Sampling Methodology This 10,000-patient sample uses stratified enrichment sampling, not a simple random draw from the full 5,000,000-patient population. A random draw at true prevalence (~8 per 100,000) would return close to zero HD-positive patients in a sample this size — useless for evaluation purposes. Instead, this sample is deliberately weighted to guarantee representation of: confirmed HD patients with full symptom trajectories, their pedigree-linked relatives (parent/child edges), the genetic biomarker panel, a realistic non-HD background population for contrast and noise modeling. The full 5,000,000-patient licensed dataset, by contrast, keeps HD at its真 real-world epidemiological rate (~0.008% / 8 per 100,000) embedded in a full-scale, non-enriched background population — matching real hospital-network case density rather than an artificially inflated one. Data Model — OMOP CDM v5.4 Standardized to OMOP CDM v5.4, maintained by OHDSI (Observational Health Data Sciences and Informatics), the current supported CDM release and the de facto standard for federated observational health research [7]. Tables included in this sample: Table Purpose PERSON Demographics: birth date, gender, race, ethnicity OBSERVATION_PERIOD Temporal coverage window per patient VISIT_OCCURRENCE Inpatient / outpatient / ER encounters CONDITION_OCCURRENCE SNOMED-CT diagnoses, staged by disease progression DRUG_EXPOSURE RxNorm-coded prescriptions MEASUREMENT Lab values and the CAG repeat genetic assay (LOINC/OMOP 4152011) FACT_RELATIONSHIP Parent-of / child-of pedigree edges Standardized Concept Dictionary (selected) Domain Concept Concept ID Vocabulary Condition Huntington's Disease 434221 SNOMED-CT Condition Chorea 376304 SNOMED-CT Condition Dysphagia 4184643 SNOMED-CT Condition Cachexia 433736 SNOMED-CT Condition Depressive Disorder 440383 SNOMED-CT Condition Injury / Trauma 432791 SNOMED-CT Condition Essential Hypertension 320128 SNOMED-CT Condition Type 2 Diabetes Mellitus 201820 SNOMED-CT Drug Tetrabenazine 10454 RxNorm Measurement CAG Repeat Length 4152011 LOINC/OMOP Relationship Parent of 40485452 OMOP Relationship Child of 41436030 OMOP Genetic Biomarker Sub-Model — CAG Repeat Assay The MEASUREMENT domain includes a CAG repeat length assay against real clinical thresholds: under 36 repeats is non-pathogenic/intermediate, 36–39 is reduced-penetrance, and 40 or more is fully penetrant [1,2]. Values in the sample span the intermediate/reduced-penetrance zone through highly-expanded, fully-penetrant alleles — consistent with testing extending across the pedigree (at-risk relatives as well as confirmed probands), not only diagnosed cases. Family Pedigree Graph FACT_RELATIONSHIP records encode a bidirectional kinship graph (parent-of / child-of edge pairs), enabling multi-generational inheritance-vector analysis and graph neural network training on hereditary transmission patterns — a structure almost never available in de-identified real-world EHR extracts. File Format Delivered as Apache Parquet (snappy compression), directly queryable with DuckDB, Polars, or PySpark without loading the full file into memory. Column-level types and schema definitions are documented in the accompanying technical file. Licensing This sample is distributed under Creative Commons Attribution 4.0 International (CC BY 4.0) — free for any use, commercial or non-commercial, with attribution. The full 5,000,000-patient dataset is a separate commercial license obtained via sentineldata.com.ua. Intended Use & Limitations This is synthetic data generated for machine learning development, algorithm validation, and educational use. It does not describe real patients and must not be used to make or support any individual clinical decision. Statistical properties are calibrated against published epidemiological and genetic literature but have not been independently peer-reviewed; users conducting downstream research should validate assumptions against primary sources cited below. Citation Sentinel Data. Huntington's Disease (HD) Enriched Synthetic Cohort — Free Sample (N=10,000), OMOP CDM v5.4. 2026. Full dataset: https://sentineldata.com.ua/dataset/huntington-s-disease-cohort-dataset References [1] American College of Medical Genetics and Genomics. Standards and Guidelines for Clinical Genetics Laboratories: Huntington Disease. Genetics in Medicine, 2014. https://www.nature.com/articles/gim2014146[2] GeneReviews (NCBI Bookshelf). Huntington Disease. NBK1305. https://www.ncbi.nlm.nih.gov/books/NBK1305/[3] Medina A, et al. Prevalence and Incidence of Huntington's Disease: An Updated Systematic Review and Meta-Analysis. Movement Disorders, 2022. PMID: 36161673.[4] Huntington Study Group. How Many People Have Huntington Disease? 2024. https://huntingtonstudygroup.org/hd-insights/how-many-people-have-huntington-disease/[5] Rare Disease Advisor. Huntington Disease Epidemiology. 2025. https://www.rarediseaseadvisor.com/disease-info-pages/huntington-disease-epidemiology/[6] Rare Disease Advisor. Addressing the Underlying Causes of Huntington Disease. 2026. https://www.rarediseaseadvisor.com/insights/addressing-underlying-causes-huntington-disease/[7] OHDSI. OMOP Common Data Model v5.4. https://github.com/OHDSI/CommonDataModel

View source

Similar papers

#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#machine learning Review Open access Jun 2014

Why Early-Stage Software Startups Fail: A Behavioral Framework

This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.

Carmine Giardino, Xiaofeng Wang, P. Abrahamsson · 175 citations · ⚡19
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#machine learning Review Open access May 2016

Key Challenges in Software Startups Across Life Cycle Stages

It is found that what perceived as biggest challenges by software startups do vary across different life cycle stages, even though its significance decreases when the learning focuses of the startups move from problem to solution and their products mature.

Xiaofeng Wang, Henry Edison, Sohaib Shahid Bajwa et al. · 62 citations · ⚡6
#computer vision Conference Aug 2008

Scrum in a Multiproject Environment: An Ethnographically-Inspired Case Study on the Adoption Challenges

Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoptio...

A. Marchenko, P. Abrahamsson · 59 citations · ⚡11

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.