Author

Yuanming Li

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Decoding cancer circulating transcriptomic signatures with language models

Current liquid biopsy methods for multi-cancer detection using plasma cell-free RNA (cfRNA, short RNA fragments circulating in blood that can reflect disease states) typically rely on gene annotations, which can overlook signals from unannotated or repetitive genomic regions. We present GeneLLM, a Transformer-based model that directly processes the nucleotide sequences of human-mapped cfRNA reads to identify cancer-indicative signatures. By bypassing gene-level quantification, the model retains signals from transcriptomic dark matter. The model learns latent pseudo-biomarkers (prototype representations from aggregated cfRNA read embeddings) that serve as discriminative features for cancer classification, rather than corresponding to explicit genomic sequences. Here we show that, in a multi-centre cohort, GeneLLM achieves ROC-AUC values ranging from 0.9250 to 0.9962 across several cancers, while maintaining comparable performance at one-sixth of the typical sequencing depth. These results suggest that sequence-level modelling of plasma cfRNA can capture diagnostically relevant information beyond annotation-dependent approaches, enabling more cost-efficient and scalable cancer screening. Cell-freeRNA (cfRNA) can be a non-invasive and cost-effective biomarker for cancer therapy and clinical outcomes, but its analysis remains challenging. Here, the authors develop GeneLLM, a cfRNA-based large language model that processes raw cfRNA data and allows accurate cancer classification from plasma biopsies.

Siwei Deng, Lei Sha, Yongcheng Jin et al. · 0 citations