This repository contains the code and analysis workflows associated with the publication "Serum Proteome Profiling Reveals a Candidate Biomarker Signature for Paediatric Chronic Nonbacterial Osteomyelitis (CNO)/Chronic Recurrent Multifocal Osteomyelitis (CRMO)” by Eve Roberts, Euan McDonnell, Phillip Brownridge, James Cassat, Claire Eyers, Hermann J Girschick, Giulia Inguscio, Ralf Knoefler, Martin W Laass, Christian Lueck, Henner Morbach, Rosemary Maher, Lauren Mee, Jenny Hawkes, Rachel Corkhill, Eva Caamano Gutierrez &, Christian M Hedrich, currently under review. This particular repository collates the in-silico workflow developed by the team at the Computational Biology Facility, LIV-SRF, University of Liverpool. The archived code includes scripts used for In silico strategy for biomarker detection. 1) Discovery cohort was split into an 80% training and 20% held-out test set. 2) Differential abundance analysis was performed on the training set to perform univariate filtration of proteins. 3) Samples were segregated into 10 cross-validation (CV) folds. 4) Least absolute shrinkage and selection operator (LASSO) was performed 100 times with 90% sub-sampling on each of the training slice of the 10 CV folds independently. 5) Candidate proteins were selected as i) consistently detected across CV Folds and LASSO repeats, log fold-change (logFC) and false discovery rate (FDR) thresholds and 4 imputation methods (Bayesian PCA (bPCA), random forest (RF), structured least-squares algorithm (SLSA) or minimum deterministic (minDet)), or unimputed (None) ii) unimputed, with LogFC > 0.5 and FDR < 0.01, or iii) correlated with a Pearson’s r > 0.7 and LogFC > 3. 6) The 10 CV-folds were filtered for the candidate proteins and each training fold used to train a random forest to predict on its corresponding test slice. 7) A model was also trained on the full 80% training data using only the candidate proteins and this used to predict on the 20% held-out test data. 8) The validation cohort data were filtered to retain only candidate proteins detected and leave-one-out CV (LOOCV) was performed to validate protein biomarker panel performance, predicting on the held-out test samples. All code and It is provided to support transparency, reproducibility and reuse of the analytical methods described in the publication. The corresponding GitHub repository contains the actively maintained version of the code: https://github.com/CBFLivUni/CNOSerumProtBiomarkers/ This piece of work is a collaborative effort primarily between the University of Liverpool (UK) and Alder Hey Children's Hospital (UK) with collaborators from The Vanderbilt University Medical Centre (USA), Vivantes Clinic Friedrichshain (Germany), Univerisity of Florence (Italy), Technische Universitat (Dresden, Germany) and University Hospital Wurzburg (Germany). Publication: currently under review, links to its fully accesible form will be updated in due course.
Euan McDonnell· Zenodo (CERN European Organi...· 0 citations
This repository contains the code and analysis workflows associated with the publication "Serum Proteome Profiling Reveals a Candidate Biomarker Signature for Paediatric Chronic Nonbacterial Osteomyelitis (CNO)/Chronic Recurrent Multifocal Osteomyelitis (CRMO)” by Eve Roberts, Euan McDonnell, Phillip Brownridge, James Cassat, Claire Eyers, Hermann J Girschick, Giulia Inguscio, Ralf Knoefler, Martin W Laass, Christian Lueck, Henner Morbach, Rosemary Maher, Lauren Mee, Jenny Hawkes, Rachel Corkhill, Eva Caamano Gutierrez &, Christian M Hedrich, currently under review. This particular repository collates the in-silico workflow developed by the team at the Computational Biology Facility, LIV-SRF, University of Liverpool. The archived code includes scripts used for In silico strategy for biomarker detection. 1) Discovery cohort was split into an 80% training and 20% held-out test set. 2) Differential abundance analysis was performed on the training set to perform univariate filtration of proteins. 3) Samples were segregated into 10 cross-validation (CV) folds. 4) Least absolute shrinkage and selection operator (LASSO) was performed 100 times with 90% sub-sampling on each of the training slice of the 10 CV folds independently. 5) Candidate proteins were selected as i) consistently detected across CV Folds and LASSO repeats, log fold-change (logFC) and false discovery rate (FDR) thresholds and 4 imputation methods (Bayesian PCA (bPCA), random forest (RF), structured least-squares algorithm (SLSA) or minimum deterministic (minDet)), or unimputed (None) ii) unimputed, with LogFC > 0.5 and FDR < 0.01, or iii) correlated with a Pearson’s r > 0.7 and LogFC > 3. 6) The 10 CV-folds were filtered for the candidate proteins and each training fold used to train a random forest to predict on its corresponding test slice. 7) A model was also trained on the full 80% training data using only the candidate proteins and this used to predict on the 20% held-out test data. 8) The validation cohort data were filtered to retain only candidate proteins detected and leave-one-out CV (LOOCV) was performed to validate protein biomarker panel performance, predicting on the held-out test samples. All code and It is provided to support transparency, reproducibility and reuse of the analytical methods described in the publication. The corresponding GitHub repository contains the actively maintained version of the code: https://github.com/CBFLivUni/CNOSerumProtBiomarkers/ This piece of work is a collaborative effort primarily between the University of Liverpool (UK) and Alder Hey Children's Hospital (UK) with collaborators from The Vanderbilt University Medical Centre (USA), Vivantes Clinic Friedrichshain (Germany), Univerisity of Florence (Italy), Technische Universitat (Dresden, Germany) and University Hospital Wurzburg (Germany). Publication: currently under review, links to its fully accesible form will be updated in due course.
Euan McDonnell· Zenodo (CERN European Organi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.