Skip to content
#protein folding Open access

CBFLivUni/CNOSerumProteinsPublic: CNO Serum Proteomics Biomarkers: peer review release, minor updates.

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This repository contains the code and analysis workflows associated with the publication "Serum Proteome Profiling Reveals a Candidate Biomarker Signature for Paediatric Chronic Nonbacterial Osteomyelitis (CNO)/Chronic Recurrent Multifocal Osteomyelitis (CRMO)” by Eve Roberts, Euan McDonnell, Phillip Brownridge, James Cassat, Claire Eyers, Hermann J Girschick, Giulia Inguscio, Ralf Knoefler, Martin W Laass, Christian Lueck, Henner Morbach, Rosemary Maher, Lauren Mee, Jenny Hawkes, Rachel Corkhill, Eva Caamano Gutierrez &, Christian M Hedrich, currently under review. This particular repository collates the in-silico workflow developed by the team at the Computational Biology Facility, LIV-SRF, University of Liverpool. The archived code includes scripts used for In silico strategy for biomarker detection. 1) Discovery cohort was split into an 80% training and 20% held-out test set. 2) Differential abundance analysis was performed on the training set to perform univariate filtration of proteins. 3) Samples were segregated into 10 cross-validation (CV) folds. 4) Least absolute shrinkage and selection operator (LASSO) was performed 100 times with 90% sub-sampling on each of the training slice of the 10 CV folds independently. 5) Candidate proteins were selected as i) consistently detected across CV Folds and LASSO repeats, log fold-change (logFC) and false discovery rate (FDR) thresholds and 4 imputation methods (Bayesian PCA (bPCA), random forest (RF), structured least-squares algorithm (SLSA) or minimum deterministic (minDet)), or unimputed (None) ii) unimputed, with LogFC > 0.5 and FDR < 0.01, or iii) correlated with a Pearson’s r > 0.7 and LogFC > 3. 6) The 10 CV-folds were filtered for the candidate proteins and each training fold used to train a random forest to predict on its corresponding test slice. 7) A model was also trained on the full 80% training data using only the candidate proteins and this used to predict on the 20% held-out test data. 8) The validation cohort data were filtered to retain only candidate proteins detected and leave-one-out CV (LOOCV) was performed to validate protein biomarker panel performance, predicting on the held-out test samples. All code and It is provided to support transparency, reproducibility and reuse of the analytical methods described in the publication. The corresponding GitHub repository contains the actively maintained version of the code: https://github.com/CBFLivUni/CNOSerumProtBiomarkers/ This piece of work is a collaborative effort primarily between the University of Liverpool (UK) and Alder Hey Children's Hospital (UK) with collaborators from The Vanderbilt University Medical Centre (USA), Vivantes Clinic Friedrichshain (Germany), Univerisity of Florence (Italy), Technische Universitat (Dresden, Germany) and University Hospital Wurzburg (Germany). Publication: currently under review, links to its fully accesible form will be updated in due course.

View source

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#protein folding Open access Sep 2026

Programmable design of functional proteins from natural language

Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or sequence constraints.

Fengyuan Dai, Shiyang You, Yudian Zhu et al. · 31 citations · ⚡3

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.