Skip to content

Evolution-Aware MSA Reasoning for Subsampling via Factor Graphs

Jul 2026 · arXiv.org · Vol abs/2607.22314 · 0 citations
Computer Science

TL;DR

Experiments show that AP-REASONER outperforms baseline subsamplers on structure-sensitive downstream tasks and enables controllable recovery of alternative protein conformations, highlighting the value of modeling MSA subsampling as a controllable optimization problem, where factor-graph reasoning offers an effective alternative to heuristic selection.

Abstract

Multiple Sequence Alignments (MSAs) provide protein language models with explicit evolutionary context, but their large depth makes subsampling unavoidable under limited token budgets. Existing strategies, including random selection, identity-based filtering, and diversity-driven sampling, are effective heuristics, yet provide limited control over the evolutionary signals retained in the subset. In this work, we recast MSA subsampling as an explicit optimization problem, where key evolutionary measures, including query identity and diversity, are treated as controllable objectives. Building on this view, we introduce AP-REASONER, an Affinity-Propagation-based factor-graph approach. With evolution-aware unary factors, exemplar-consistency factors, and two control knobs, AP-REASONER performs factor-graph reasoning through message passing to infer a fixed-budget MSA subset. Experiments on long-range contact prediction and conformational ensemble prediction show that AP-REASONER outperforms baseline subsamplers on structure-sensitive downstream tasks and enables controllable recovery of alternative protein conformations. These results highlight the value of modeling MSA subsampling as a controllable optimization problem, where factor-graph reasoning offers an effective alternative to heuristic selection.

View source

Similar papers

Preprint Jul 2026

OptGraph: Large Language Models Enhanced Evolutionary Optimization Via Graph Retrieval-Augmented Generation

OptGraph is the first optimization agentic workflow that introduces graph retrieval-augmented generation (GraphRAG) and first constructs reusable experience as a typed graph, capturing the relationships among modeling patterns, problem formalization, implementation details, and error corrections.

Xianchao Xiu, Jianhao Li, Huangyue Chen et al. · 1 citation
Jul 2026

Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation

This work proposes a framework that ensembles the reasoning structure, not just the answers, of multiple LLMs by weighted merging of Directed Acyclic Graphs (DAGs) extracted from reasoning chains by weighted merging of Directed Acyclic Graphs (DAGs) extracted from reasoning chains.

Amruta Parulekar, Jinu Lee, Dilek Z. Hakkani-Tür et al. · 1 citation
Open access Aug 2026

Pruning the Search, Not the Signal: Adaptive-Banding Needleman–Wunsch via Protein Language Model Confidence

Dynamic programming (DP) yields exact quadratic-time (O(NM)) pairwise sequence alignments. Static banding heuristics (O(NW)) fail catastrophically on low-identity (< 30%), asymmetric insertions/deletions (indels), or extreme length ratios, dropping core-block Sum-of-Pairs (SP) score recovery to 20%– 50%. Conversely, recent protein language model (PLM) aligners evaluate all N × M cells without search grid constraints. To bridge this gap, we introduce Adaptive-Banding Needleman–Wunsch (AB-NW), leveraging PLM contextual representations to construct a confidence-adaptive DP corridor prior to fine-resolution DP while keeping downstream scoring unmodified. AB-NW downsamples residue embeddings, computes a coarse alignment, and sets per-row corridor bounds via normalized confidence metrics. Evaluated via JIT-compiled buffers, this reduces time complexity to and space to O(NWmax) . Benchmarked across three PLM backbones (ESM2-8M, ESM2-35M, ProtBERT) across nine structural challenge categories, AB-NW recovers > 98.9% of exact unconstrained alignment scores and core-block SP accuracy across static banding failure modes (Twilight Zone, Asymmetric Indels, Extreme Aspect Ratios) while eliminating 55.3%–78.8% of active DP cells. On large protein matrices (N, M ≥ 3, 700), AB-NW eliminates 87.6%– 91.7% of cells, achieving speedups of 9.79×–13.30× (pure DP) and 1.73×– 2.94× (end-to-end), reaching up to 18.12× on unbiased controls (p < 0.05 to p < 10−15), making AB-NW practical for large-scale, high-throughput sequence alignment pipelines.

Muhammad Shoaib, Waqas Ali · 0 citations
Preprint Aug 2026

LLM-Guided Graph Generation for Structure-Based Local Improvement Methods

An automatic pipeline that is problem-agnostic to all problems in the MiniZinc format is built, finding that algorithm selection achieves a 39.6% average problem-weighted win rate against a one-shot Gurobi baseline, more than doubling the best single configuration (19.3%).

Hai Xia, Vaidyanathan Peruvemba Ramaswamy, Stefan Szeider · 0 citations
#natural language process... Preprint Aug 2026

QUORUM: QUality-Optimized Routing Using Multiple annotators

This work introduces QUORUM (QUality-Optimized Routing Using Multiple annotators), a budget-aware routing framework that dynamically assigns each instance to human or LLM annotators under a fixed annotation budget and supports multiple annotations per instance, combining them through agreement-based rewards to improve reliability.

Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu et al. · 0 citations
Conference 2026

Learning Unified Graph and Language Representations for SMT Algorithm Selection

Evaluated across nine SMT logics, SMT-Select consistently outperforms existing selectors and SMT-COMP winning solvers and closes at least 30% of the performance gap between the competition winner and the virtual best solver (VBS).

Zhengyang Lu, Paul Sarnighausen-Cahn, Jiahao Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.