Evaluating Autonomous LLM Agents Across Molecular Prediction and Optimization Benchmarks
Large language model (LLM) agents are increasingly capable of carrying out autonomous computational research, but it remains unclear whether they can develop molecular modeling methods that compete with strong human-developed approaches. Here, we evaluate autonomous method development across four settings: Therapeutics...