The rapid evolution of generative models has unlocked new potentials in protein binder design, a pivotal task in structural biology, by facilitating end-to-end generation via joint sequence-structure modeling or hallucination. However, existing approaches are predominantly implemented under a single-target, single-state assumption, limiting their ability to model multi-target or multi-state interactions required for advanced function-oriented protein design. Here, we introduce Chamaileon, which unifies multi-target and multi-state binder design by formulating the problem as cross-context binding landscape modeling. The framework is underpinned by a training paradigm termed In-Context Complex Co-Design (I3CD) for context-aware sequence-structure co-modeling. During inference, we employ Mixture-of-Paths Sampling (MoPS), a scalable strategy that optimizes a single sequence across contexts while alleviating the scarcity of high-quality multi-conformational paired data. Extensive evaluation on our newly constructed benchmark, CROSS, demonstrates that Chamaileon effectively generates sequences adaptable to diverse conformational landscapes and multi-target requirements. The code is available on https://github.com/caohengyuan/Chamaileon.
This paper proposes UniEdit, a Unified Graph-based Mixture-of-Experts (MoE) Molecular Editing model that offers a robust alternative to LLMs and incorporates a Mixture-of-Experts architecture that dynamically routes tasks to specialized components.
Jiajun Yu, Zhihao Wu, Yizhen Zheng et al.· Proceedings of the 32nd ACM...· 0 citations
MCTH (Monte Carlo Tree Hallucination), an inference-only framework that casts all-atom sequence-structure co-design as uncertainty-aware planning over hallucinated states from pretrained folding and inverse-folding models, with optional biophysical control within the same decision loop is introduced.
Xuefeng Liu, Mingxuan Cao, Xiao Luo et al.· 0 citations
Balancing target-specific biological affinity with drug-likeness remains a central challenge in de novo molecular design, where existing generative models often exhibit limited controllability or reduced structural diversity. Here, we present DF-S4, a conditional molecular generation framework based on Structured State Space Models (S4), which addresses this limitation through a disentangled latent representation and hierarchical feature-wise linear modulation (FiLM). By decoupling structural and property variables and injecting conditional signals across multiple representation levels, DF-S4 enables fine-grained and stable multi-objective control beyond conventional input-level conditioning. Evaluated via a rigorous progressive multi-denominator auditing framework across three kinase targets (EGFR, BRAF, and FGFR1), DF-S4 exhibits robust target-steering performance, yielding favorable intradomain active ratios (74.2-83.9%) while maintaining high novelty (>95%) and competitive internal diversity (∼0.85). Furthermore, DF-S4 shifts the multi-objective Pareto Frontier toward regions of simultaneous high apparent affinity and drug-likeness, mitigating the distributional-trapping trade-offs commonly observed in prior approaches. Molecular docking analyses confirm the physical plausibility of generated candidates, yielding stronger binding affinities than reference inhibitors. Finally, comprehensive ablation studies and latent space dependence analyses mathematically validate that both explicit latent disentanglement and hierarchical FiLM modulation are critical for robust feature isolation, tighter property alignment, and generative stability.
Yuecheng Peng, Yongquan Jiang, Baoxue Quan et al.· Journal of Chemical Informat...· 0 citations
Abstract Motivation To enable real-world protein-ligand affinity prediction, not only out-of-distribution generalization but also robustness to variable structural availability and quality should be considered in model design. Results We present AlignNet, a hierarchical representation alignment framework that mitigates intra- and inter-molecular heterogeneity to learn robust protein-ligand embeddings for generalizable affinity prediction, even from sequence-level inputs. Its intra-molecular module projects unimodal and multimodal features into a unified space, aligning augmented multimodal views for feature fusion and unimodal with multimodal embeddings to distill multimodal priors for structure-agnostic inference. Its inter-molecular module aligns protein and ligand embeddings for cross-molecular integration. Extensive experiments show that AlignNet (i) achieves highly competitive performance, with up to a 20.4% gain in SCC on the challenging LBA 30% split under sequence-only settings, suggesting improved out-of-distribution generalization; and (ii) learns well-separated affinity-related clusters, supporting reliable structure-independent prediction. Availability and implementation AlignNet is available at https://github.com/altriavin/AlignNet.
Xiaowen Hu, Hongyi Huang, Hao Sun et al.· Bioinformatics· 0 citations
Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we introduce Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions. We compile CREED-CCV-2+USPTO-XL, an ultra-large-scale dataset of ~45.6 million verified reactions to train the C3LM (Chemistry Constraint-Consistent Language Model). By integrating fine-tuning with ChemCensor-based and novelty-oriented rewards, our model achieves state-of-the-art performance on the OOD URSA-expert-2026 benchmark. Further analysis of reaction uniqueness shows that LLMs and conventional models explore complementary reaction spaces, motivating ensemble-based retrosynthesis systems. Overall, our results establish Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.
B. Zagribelnyy, Ivan D. Ilin, N. Bondarev et al.· 0 citations