Skip to content

Integrating Plasma Proteomics and Polygenic Risk Scores for the Prediction of Lung Cancer Risk.

Sep 2026 · International Journal of Cancer · 0 citations · 23 references
Medicine

Abstract

Multimodal diagnostic strategies and risk stratification are essential for improving early lung cancer detection, clinical outcomes, and personalized interventions. To develop a predictive model for lung cancer incidence by integrating plasma proteomics, genetic factors, and demographic data. We analyzed data from the UK Biobank, which included 47,655 adults without a history of lung cancer at baseline, with 641 incident cases and a median follow-up duration of 13.87 years. COX proportional hazards modeling with SHapley Additive exPlanations (SHAP) sequential feature selection was used to identify significant plasma proteins associated with lung cancer. An XGBoost-based survival model was constructed and evaluated to assess its predictive utility. Proteomic risk score (ProRS), demographic, and Combined risk scores were developed to evaluate their predictive capabilities across genetic risk strata. Among 2911 plasma proteins, we identified 15 significant plasma proteins associated with the development of lung cancer. The XGBoost survival model demonstrated promising predictive performance, achieving a C-statistics of 0.803. Risk stratification revealed that high-risk groups had up to a 101-fold increased risk of lung cancer compared to low-risk groups. We analyzed expression patterns of 15 carcinogenesis-associated proteins during the 15-year period preceding diagnosis. TNR and CLEC3B were identified as the earliest dysregulated proteins, with expression levels consistently downregulated beginning as early as 14 years before diagnosis. GDF15 exhibited sustained and pronounced upregulation throughout the entire 14-year observation window. Integrating plasma proteomics, polygenic risk scores, and demographic factors improves personalized lung cancer risk stratification, offering potential for refined screening strategies.

View source

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#protein folding Open access Sep 2026

Programmable design of functional proteins from natural language

Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or sequence constraints.

Fengyuan Dai, Shiyang You, Yudian Zhu et al. · 31 citations · ⚡3

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.