The Algorithm Question: When Architectural Complexity Meets Biological Reality in Antimicrobial Peptide Prediction
Abstract
Computational prediction and design of antimicrobial peptides (AMPs) has progressed quickly, moving from physicochemical descriptors paired with classical machine learning to transformer architectures and generative models. Each new architecture claims improved performance, yet basic questions remain unanswered. Does added complexity produce proportionally better biological predictions, and what do different architectures actually learn? Recent panoramic surveys have cataloged progress; the present review takes a different stance, arguing that the field has favored architectural novelty over biological understanding and has widened the gap between computational metrics and therapeutic relevance. Three findings emerge from a re‐examination of the evidence. First, deep learning models do not consistently outperform well‐designed shallow approaches, and both encode statistically similar chemical information. Second, reported in vitro validation rates of 59% to 100% for heavily pre‐screened designs cluster within a narrow band irrespective of algorithm class, which suggests that pre‐screening pipelines, rather than raw model accuracy, drive these numbers. Third, generative models such as HydrAMP, Multi‐CGAN, and latent‐diffusion frameworks sample within the physicochemically constrained envelope defined by their training data, and because antimicrobial activity can change sharply with a single substitution, descriptor‐space novelty does not by itself establish functional novelty. To support such a critique with original evidence, the validation‐rate analysis is extended across 18 representative studies. Two community instruments are then proposed: a Biology‐First Algorithm Selection (BFAS) decision framework that operationalizes a four‐axis selection protocol over task, data, interpretability, and compute budget; and AMP‐CASP, an AMP‐equivalent of the CASP blind‐benchmark that turns an earlier call for a CASP‐like community assessment into an operational protocol, with held‐out activity, toxicity, and stability labels, homology‐aware splitting, organism‐resolved endpoints, and explicit assay‐condition metadata. Progress now depends on an honest assessment of what algorithms can and cannot deliver in moving from computational prediction to clinical translation.