Author

A. Carbone

2 papers indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

X-PAIR: an ultrafast multitask framework for proteome-scale reconstruction of PPI networks and partner-specific interfaces from sequence

Protein–protein interaction prediction and residue-level interface localisation are biologically intertwined but usually treated as separate computational problems. Here we present X-PAIR, a sequence-based multitask deep learning framework that jointly predicts whether two proteins interact and identifies their partner-specific interface residues. By combining protein language-model representations with lightweight cross-attention, X-PAIR requires neither structural templates nor multiple-sequence alignments. Across leakage-controlled benchmarks, it outperforms existing methods in both tasks, with substantial gains in interface localisation. Multitask learning preserves single-task performance while returning both outputs at near-single-task cost, enabling one million protein pairs to be analysed in under two hours—approximately 500-fold faster for interface prediction and 20-fold faster for PPI prediction than current approaches—thereby enabling proteome-scale analysis. Cross-species analyses reveal distinct evolutionary dependencies: interaction prediction benefits from multispecies training, whereas interface localisation remains robust across taxonomic scales. X-PAIR thus links proteome-scale interaction discovery to the residue-level determinants of partner-specific molecular recognition.

S. Rescalli, A. Carbone · 0 citations
Open access Aug 2026

PLMView: collaborative protein language model representations for fast and scalable specialized protein function inference

The functional classification of protein sequences remains a major bottleneck in biology. Although protein language model (PLM)-based approaches have substantially improved broad protein function prediction, most protein sequences still lack precise annotation at the level of specialized functions—the fine-grained molecular roles that define specificity within protein families. We present PLMView, an unsupervised framework for fine-grained protein function classification directly from sequence. PLMView reframes protein function inference as a relational problem: instead of embedding sequences in isolation, it positions them within a collaborative functional space defined by comparisons with PLM embeddings of anchor sequences, thereby capturing subtle sequence–function relationships. Without requiring labeled data, family-specific training, or PLM fine-tuning, PLMView accurately distinguishes specialized functions among homologous proteins and highlights residues likely to determine functional specificity. The method achieves high precision while remaining computationally efficient, classifying approximately 10,000 sequences with 1,000 anchors in under 40 minutes; compared with pooled-embedding approaches and, in challenging cases, Sequence Similarity Networks, PLMView provides finer and more biologically coherent functional resolution, while achieving more than 10-fold speed-up over SSN reconstruction on datasets of this scale. Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution and transcriptional deployment across the global ocean.

Vinh-Son Pho, Alessandro Natale Bianchi, Mattéo Scarsini et al. · 0 citations