Benchmarking Antibody Modeling Tools across Structure Prediction, Docking, and Paratope–Epitope Interface Analysis
Computational antibody engineering requires reliable prediction of antibody variable-fragment structures, antigen–antibody complexes, and binding interfaces. However, publicly available tools for these tasks have rarely been compared across the complete workflow under a controlled and statistically grounded design. We evaluated ImmuneBuilder, IgFold, AlphaFold3, GRAMM, and dyMEAN on 50 non-redundant humanized antibody–antigen complexes using multiple retained predictions and paired statistical testing. All three antibody structure predictors were accurate, with AlphaFold3 performing best overall and for the third complementarity-determining region of the heavy chain. AlphaFold3 also substantially outperformed GRAMM and dyMEAN in complex prediction, producing medium- or high-quality binding interfaces for 46% of the complexes, although overall interface accuracy remained limited. When docking was reliable, AlphaFold3 accurately recovered epitope and paratope residues, salt bridges, and non-bonded contacts, but reproduced hydrogen bonds and fine-grained contact strengths less consistently. These findings provide practical guidance for selecting tools across antibody-modeling workflows and identify persistent limitations in fine-grained interface prediction. Data, structural predictions, evaluation results, and analysis code are available from Zenodo under record 20710876.