Deep generative models have become a leading approach for designing therapeutic molecules, yet efficiently exploring vast biomolecular sequence spaces remains difficult, particularly for targets with limited training data. The prior distribution that seeds a generative model shapes which regions of sequence space it explores, and recent work suggests that non-classical distributions sampled from quantum processors can serve as a structured alternative to the factorised Gaussian priors used by default. Whether such priors help on complex biological design tasks has been largely untested. Here we present what is, to our knowledge, the first end-to-end hybrid quantum-classical pipeline for de novo design of MHC class I-binding peptides, coupling a generative adversarial network (GAN) to latent vectors sampled from a real photonic quantum processor. Tested in silico across 131 HLA alleles, quantum-derived priors increased the yield of predicted strong binders, with the largest relative gains for understudied alleles where classical baselines perform worst. We selected three understudied alleles for further evaluation, finding that large gains coincided with broader sequence exploration at non-anchor positions while anchor specificity was preserved. On these three alleles, we validated the designs in vitro using peptide-MHC stability ELISAs, confirming that quantum-designed peptides are potent stabilisers of peptide-MHC class I complexes. These results establish structured, hardware-realisable non-classical priors as a useful inductive bias for generative peptide design, with direct relevance to personalised immunotherapies and vaccines.
How advances in artificial intelligence and computational modeling may reshape the rational design of next-generation peptide therapeutics is explored and an integrated experimental–computational framework is proposed to facilitate the development of clinically actionable candidates is proposed.
Ha Thi Ngoc Nguyen, B. Le, Nhung Thi Hong Van et al.· Pharmaceuticals· 0 citations
Parameterized quantum circuits (PQCs) provide a flexible substrate for hybrid quantum machine learning (QML), but their practical value on Noisy Intermediate-Scale Quantum (NISQ) devices remains an empirical question, especially because training depth and scale can introduce optimization challenges such as barren plateaus. Here we study how the number and topology of two-qubit entangling gates in the feature-map stage influence a fixed hybrid QNN workflow for classifying strong versus weak epitope-receptor binding in Porcine Reproductive and Respiratory Syndrome (PRRS) vaccine design. The dataset consists of docking-derived binding affinities for N=80 9-mer epitopes, labeled as Strong or Weak binding, and partitioned into training, validation, and test subsets using a 40:30:30 split. We compare a classical CNN benchmark with a hybrid Embedding-QNN architecture under four feature-map configurations: a non-entangling Z feature map, an all-to-all high-entanglement ZZ feature map, and two interleaved nearest-neighbour entanglement patterns of low and high depth. Among the configurations tested, the high-entanglement ZZ feature map is seen to provide the strongest evidence of reduced training-set overfit, with a lower training area under the accuracy curve (AUAC) and the highest test/training AUAC ratio, while preserving competitive test-set accuracy. These results do not establish a general QML advantage, but they suggest that feature-map entanglement topology is a meaningful design variable for sparse biological screening tasks and warrants further evaluation with additional metrics, larger datasets, and noise-aware or hardware-based experiments.
Aspen Erlandsson Brisebois, Luis Pablo Gonzalez Dominguez, Shivansi Prajapati et al.· arXiv.org· 0 citations
The rapid growth of quantum computing is driven by promises of performing complex calculations with unprecedented speed; however, current use cases have been limited by quantum hardware and the difficulty of identifying problems that classical computers cannot easily address. Within these constraints, biological problems including drug discovery, protein folding, and precision medicine present an opportunity to understand how current quantum hardware can make advances. In immunology, accurate prediction of cancer neoantigens remains a major challenge, limited by small, noisy datasets and the inability of classical models to generalize. In approaching the problem, we explore multiple noise mitigation techniques, including Pauli twirling and dynamical decoupling, in conjunction with controlled shot-based sampling to stabilize training on real hardware and in a warm start hybrid approach. With these approaches, we demonstrate the use of Quantum Convolutional Neural Networks (QCNNs) for both MHC binding and immunogenicity prediction, including a quantum hardware experiment involving 46 qubits that achieved a 6% increase in classification accuracy with fewer training samples compared to classical approaches. Building on these models, we introduce Quantum Convolutional HLA Immunogenic Peptide Prediction (Q-CHIPP), a combinatorial framework integrating MHC binding and T-cell recognition. It targets HLA-A*02:01–restricted 9-mer peptides and, more accurately, identifies those peptides known to be immunogenic, improving the prognostic impact of predicted neoantigen load. Together, these represent a large-scale application of QCNNs in biomedical modeling, highlighting both the feasibility and promise of quantum machine learning for data-limited biological systems and establishing a scalable foundation for quantum-enhanced biomedical research.
Ryan Peters, Kahn Rhrissorrakrai, Prerana Bangalore Parthasarathy et al.· Science Advances· 0 citations
Virtual screening (VS) is an essential tool in drug discovery to prioritize potential drug candidates from vast chemical space. One key challenge limiting its performance is accounting for protein conformational flexibility. While ensemble docking methods have been developed to address this challenge by incorporating multiple protein conformations, these methods often rely on computationally intensive physics-based simulations to sample the relevant conformational space. Generative machine learning models offer a highly promising, scalable, and high-throughput alternative to overcome the limitations of these traditional approaches. We therefore investigate whether conformational ensembles generated by BioEmu, a recently developed generative model, can improve VS performance for kinase targets. Using the DUD-E benchmark data set and a validated AutoDock-GPU protocol, we generated and analyzed nearly 1300 structures across 26 kinases (approximately 50 structures each). BioEmu produces structurally diverse ensembles with substantial performance variation among individual structures. However, ensemble methods employing consensus or best-score selection fail to improve upon, and often degrade, VS performance compared to crystal structure baselines. To investigate the source of this limitation, we quantified the relationship between KinCoRe-based conformational state classification and screening performance. By calculating the coefficient of determination (R2) across the kinase subset, we found that the structural features governing VS performance differ substantially from those defining standard conformational states, with KinCoRe classifications leaving over 84% of performance variance unexplained. This critical gap demonstrates that structural diversity alone is insufficient to guarantee screening success. We show that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.
Jaeoh Shin, K. Joo, Jejoong Yoo· Journal of Chemical Informat...· 0 citations
An unsupervised machine-learning framework that leverages hybrid high-dimensional peptide representations to discover high-performance AFPT families without requiring 3D structures or large labeled data sets is presented and demonstrates how unsupervised hybrid-feature learning can reveal actionable biophysical design rules from sequence data alone.
Nazmul Shuzan, Jialun Wei, Jie Zheng· Journal of Chemical Informat...· 0 citations
Protein engineering is limited less by generating variants than by the cost of evaluating them, so designing under a tight budget demands sequence features that let a model learn fitness from very few examples. We introduce ALSEBO (Active Learning Sequence Exploration via Bayesian Optimization), which couples a generative latent sequence landscape to Bayesian optimization and featurizes candidates with direct-coupling-analysis (DCA) coevolutionary statistics. This representation carries a specific inductive bias: it places the dominant organizer of the fitness landscape along a single linear coordinate, producing a smooth, funnel-like objective that a low-data surrogate navigates efficiently. On a virtual avGFP fluorescence benchmark, ALSEBO reaches the optimum in ∼40 evaluations and outpaces protein-language-model embeddings and raw latent coordinates; controls with representation-neutral oracles confirm that the advantage is intrinsic, not an artifact of the benchmark. Molecular dynamics of the optimized variant recovers structural hallmarks of fluorescence, and ALSEBO transfers to divergent GFP orthologs and to a non-GFP enzyme, establishing a data-efficient route to protein design.
D. P. Kulathunga, Divyanshu Shukla, D. Potoyan· bioRxiv· 0 citations