PRIMAD-LID in Practice : An RO-Crate Profile for Reproducible ML in Bioinformatics
Abstract
Machine learning reproducibility requires documentation of interdependent research components and the decisions made throughout a study. This poster presents ongoing work to develop an RO-Crate profile for representing reproducibility metadata in machine learning–based bioinformatics studies, guided by PRIMAD-LID and illustrated through scVI. PRIMAD-LID (Aloqalaa et al., 2026) is a discipline-diagnostic framework for computational reproducibility that retains the six PRIMAD dimensions (Freire et al., 2016): Platform, Research Objective, Implementation, Method, Actor and Data. Three cross-cutting modifiers qualify each dimension: Lifespan captures temporal considerations; Interpretation records contextual information and rationale; and Depth specifies the required level of detail and granularity. To establish the metadata foundation, source elements from ten selected reproducibility resources spanning multiple disciplines were mapped to PRIMAD-LID (Aloqalaa et al., 2026). Candidate metadata attributes were identified through direct extraction or analytical derivation and harmonised across disciplinary contexts. The proposed profile builds on this attribute set, using RO-Crate, a research packaging approach(Soiland-Reyes et al., 2022), to represent research arteficts and contextual metadata through entities, properties and relationships in a machine-readable form with human-readable views. An illustrative application to single-cell RNA sequencing (scRNA-seq) analysis using scVI (Lopez et al., 2018) targets data provenance, preprocessing and modelling choices, software dependencies, training parameters and outputs. References: Aloqalaa, M., Soiland-Reyes, S., & Goble, C. (2026). PRIMAD-LID: A developed framework for computational reproducibility [Preprint]. arXiv. DOI: 10.48550/arXiv.2601.02349. Freire, J., Fuhr, N., & Rauber, A. (Eds.). (2016). Reproducibility of data-oriented experiments in e-Science (Dagstuhl Seminar 16041). Dagstuhl Reports, 6(1), 108–159. DOI: 10.4230/DagRep.6.1.108. Gundersen, O. E., Cappelen, O., Mølnå, M., & Nilsen, N. G. (2025). The unreasonable effectiveness of open science in AI: A replication study. Proceedings of the AAAI Conference on Artificial Intelligence, 39(25), 26211–26219. DOI: 10.1609/aaai.v39i25.34818. Lopez, R., Regier, J., Cole, M. B., Jordan, M. I., & Yosef, N. (2018). Deep generative modeling for single-cell transcriptomics. Nature Methods, 15(12), 1053–1058. DOI: 10.1038/s41592-018-0229-2. Osborne, C., Ding, J., & Kirk, H. R. (2024). The AI community building the future? A quantitative analysis of development activity on Hugging Face Hub. Journal of Computational Social Science, 7(2), 2067–2105. DOI: 10.1007/s42001-024-00300-8. Sajadieh, S., et al. (2026). Artificial Intelligence Index Report 2026. Stanford University. DOI: 10.48550/arXiv.2606.15708. Soiland-Reyes, S., et al. (2022). Packaging research artefacts with RO-Crate. Data Science, 5(2), 97–138. DOI: 10.3233/DS-210053.