Preprint
Aug 2026
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
DSAgentBench is introduced, the first benchmark to evaluate whether agents can automate full data-science workflows inside real computer environments, and reveals a substantial capability gap between current agentic systems and real data-science workflows.
Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub et al.
· 0 citations