Skip to content

Scaling Scientific Discovery Environments for Turn-Level Agentic RL

Jul 2026 · arXiv.org · Vol abs/2607.28990 · 2 citations · 56 references
Computer Science

TL;DR

Experiments show that SciDisco-14B reaches state-of-the-art on hypothesis-driven scientific data analysis benchmarks, and SciDisco-14B reaches state-of-the-art on hypothesis-driven scientific data analysis benchmarks.

Abstract

Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution environment and produces a statistical claim. Long-horizon scientific analysis remains constrained by the lack of process supervised environments over real-world scientific data. This paper introduces SciDisco, a scalable framework for training Scientific Discovery agents in process-verifiable environments. SciTh\`eque compiles hypotheses, datasets, hidden evidence graphs, and verifiers into task environments where analytical progress can be checked during interaction. DAG-grounded trajectory synthesis uses these environments to construct verifier-filtered multi-turn demonstrations. DiscoPO then uses the environment as the source of training signal, assigning turn-level credit to actions that produce verifiable analytical evidence. Experiments show that SciDisco-14B reaches state-of-the-art on hypothesis-driven scientific data analysis benchmarks.

View source

Similar papers

Preprint Aug 2026

The Little Scientist: LLM Agent-Driven Discovery via the Scientific Method

The Little Scientist is presented, a framework in which a LLM agent stepping through the scientific method can discover both new algorithms and new ensemble strategies that outperform prior solutions, and is demonstrated on two problems that require fundamentally different modes of discovery.

T. Smith · 0 citations
Review Open access Jul 2026

LLM-Powered Agentic Data Science: Automated Analysis and Insight Generation

It is argued that verification, not generation, is the binding constraint for trustworthy automated analysis in agentic data science: systems in which an LLM coordinates exploratory analysis, query generation, hypothesis formation, and reporting with limited human supervision.

M. Keerthika · 0 citations
Preprint Aug 2026

ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond

Results on diverse long-horizon benchmarks demonstrate the efficacy of ScienceFlow's ability to sustain effective research processes, and demonstrates that efficient state management, adaptive exploration, and objective-aligned execution are critical for scaling autonomous research beyond short-horizon interactions.

Ming-Ming Zhao, Jiqian Dong, Kangping Xu et al. · 0 citations
Preprint Aug 2026

AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale

This work introduces AgentMercury, a scalable framework for synthesizing executable environments from high-level business scenarios that improves substantially on both enterprise workflows and out-of-domain benchmarks spanning reasoning, coding, scientific computing, and tool use.

Minbyul Jeong, Chanwoong Yoon · 0 citations
Preprint Aug 2026

Prime Agent: A Self-Improving RLM Harness

Low-friction, expressive membrane prevents harness failures from becoming model failures and pushes measurement toward the model's true maximal underlying capability, on Factorio, where refinement allows for continuous technology progression and dedicated subagents enable parallelized work.

Seth Karten, Alex L. Zhang, Kevin Thomas et al. · 2 citations · ⚡1
Review Aug 2026

Intern-S2-Preview: Scientific Agentic Foundation Model

Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings.

Lei Bai, Jiaqi Cao, Chiyu Chen et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.