Skip to content
Preprint

Scientific Data Skills: Enabling Agent-Ready Scientific Data Services at Scale

Aug 2026 · 0 citations · 52 references
Computer Science

TL;DR

Results indicate that SciDSK improves how agents locate and understand scientific datasets, providing a stronger foundation for actionable scientific data use.

Abstract

Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for reliable dataset discovery and interpretation, constraining their effective use in scientific workflows. This limitation arises because agents must search across heterogeneous repositories and reconstruct dataset-specific semantics and operating procedures from documentation designed primarily for human use. To address this limitation, we introduce the Scientific Data Skill (SciDSK), an agent-ready representation that packages dataset-specific knowledge and operational guidance as a reusable agent skill. A SciDSK integrates dataset descriptions, scientific context, file organization, task-specific usage procedures, quality checks, and provenance information while retaining the underlying data in its original repository. We define a structured SciDSK specification and develop a systematic construction pipeline that grounds each SciDSK in authoritative dataset records and associated supporting materials. We further establish the Scientific Data Skill Bank, a unified platform that publishes SciDSK resources across six scientific disciplines and supports package access, persistent identification, and traceability to source datasets. We evaluate SciDSK through a retrieval benchmark for dataset discovery and controlled cases for dataset interpretation. On the query retrieval benchmark, Agent-SciDSK achieves 80.77% Hit@1, exceeding Agent-Raw by 9.62 percentage points. Across controlled interpretation cases, the SciDSK condition satisfies 23 of 24 assessment criteria, compared with 22 under the web-page condition. These results indicate that SciDSK improves how agents locate and understand scientific datasets, providing a stronger foundation for actionable scientific data use.

View source

Similar papers

Book Open access Jul 2026

Toward a Trustworthy and Accessible Scientific Data Workflow Platform with StreamCI

Scientific research workflows increasingly involve not only structured streaming data but also raw artifacts and derived products, requiring platforms that provide trustworthy data protection and accessible interfaces beyond simple ingestion and storage. We present extensions to StreamCI, a cloud-based streaming data management platform, that evolve it into an end-to-end scientific data workflow platform. Key advancements include expanded data lifecycle support (blob ingestion, raw data refinement, and derived data generation), a migration from RabbitMQ to Apache Kafka for scalable dataflow infrastructure, automated backup mechanisms for data reliability, and a researcher-friendly web portal paired with a Python client API library (PyStreamCI). We demonstrate these capabilities through end-to-end workflows from active research use cases and report on operational experience following deployment on Purdue University’s Anvil cloud infrastructure. In production, these extensions enable researchers across domains—such as building energy sustainability, pavement condition monitoring, and precision audiology—to manage complete data workflows, from raw artifacts to analysis-ready products, without requiring backend expertise.

Jaewoo Shin, Mehak Jain, S. Jha et al. · 0 citations
Jul 2026

SciDataSailor: Deep Scientific Data Exploring

This work presents SciDataSailor, a framework for synthesizing tool-interactive trajectories by balancing broad exploration with targeted exploitation and presents SciDataSailor, a framework for synthesizing tool-interactive trajectories as Monte Carlo Tree Search (MCTS) with four task-specific mechanisms.

J. Rao, Yicheng Qiu, Chi Zhang et al. · 1 citation
Preprint Aug 2026

Multi-Agent Discovery and Resource-Aware Autonomous Exploration of Scientific Datasets

Modern scientific facilities and instruments generate datasets at scales that are difficult for individual researchers to discover, access, and explore. Although many datasets are publicly available, using them often requires familiarity with repository organization, data formats, multiresolution structures, and visualization parameters. We present WebVisus, a constrained and resource-aware multi-agent system for discovering and autonomously exploring remote, multiresolution scientific datasets. Given a natural-language research question, WebVisus identifies the user's intent and launches an autonomous exploration agent that examines slices, volumes, and timesteps while adapting data resolution and retrieval quality to available client memory and computational resources. This design supports progressive exploration without complete dataset downloads or manual configuration of low-level visualization parameters using natural languages. We report the system architecture, constrained agent protocol, resource-aware access mechanism, and case studies evaluating autonomous visual exploration and resource-aware agentic access across scientific datasets.

Aashish Panta, Hugo Lee, G. Scorzelli et al. · 0 citations
#large language models Open access Sep 2026

From Data Quality to Quality of Agentic Data Use: A Conceptual Framework for Agentic Data Engineering

Large language models and AI agents are extending data-engineering automation beyond isolated artifact generation toward end-to-end processes in which agents interpret requirements, select data, generate transformations, invoke tools, validate results, and communicate analytical outputs. This shift introduces risks that conventional notions of data quality and execution success do not fully capture. A dataset may satisfy established quality standards, and a generated query may execute without technical errors, while the agent still selects an incorrect metric, combines incompatible analytical grains, accesses unauthorized data, or draws conclusions that are insufficiently supported by evidence. This paper develops a conceptual framework for Agentic Data Engineering centered on Quality of Agentic Data Use, defined as the extent to which an agent uses and communicates data in accordance with task, semantic, quality, security, governance, and provenance requirements. An evidence-informed analysis of Data Contracts, Semantic Layers, Data Quality, Guardrails, AI Governance, and Data Provenance shows that these foundations provide essential but fragmented capabilities. The proposed framework integrates and extends them through four core artifacts: Agentic Data Contracts, Agentic Expectations, Agentic Data Provenance, and Agentic Data Governance. It also introduces an execution lifecycle, a reference architecture, a failure taxonomy, and a multidimensional evaluation framework. A governed sales-analysis scenario illustrates how the proposed artifacts interact throughout an agent-mediated data process. In addition, a controlled Databricks prototype and a complementary benchmark comprising 10 cases and 40 executions demonstrate the framework’s technical feasibility and support the independent computation of enforcement indicators. The benchmark highlights the value of separating generation from validation while also showing that the current validation and automated-repair mechanisms require further calibration. These preliminary findings do not establish generalized improvements in safety, correctness, or reliability. Rather, they provide an operational foundation for broader empirical evaluation of trustworthy agent-mediated data-engineering processes.

Ania Cravero, Jorge Díaz, Zihao Xiao · 0 citations
Review Open access Jul 2026

FRED enables standardized FAIR metadata generation and management for omics research

Scientific research relies on transparent dissemination of data and its associated interpretations, including raw data, metadata, experimental design, and data processing details. Production and handling of research data represents an ongoing challenge, extending beyond publication into individual facilities, institutes and research groups, often termed Research Data Management (RDM). It is foundational to scientific discovery and aligned with the FAIR principles. Although the majority of peer-reviewed journals require raw data deposition in public repositories in alignment with FAIR principles, metadata frequently lacks standardization, hindering effective utilization and sharing of research findings. Here we present FRED, a generalized toolkit for FAIR metadata management in omics research based on a flexible, machine-readable YAML format. FRED enables (i) guided, dialog-based creation of metadata files, (ii) structured semantic validation, (iii) logical cross-file search, (iv) API-based integration with external systems, and (v) self-hosted web deployment. We demonstrate the utility of FRED through a complete annotation workflow applied to a published single-nucleus RNA-seq dataset, covering metadata generation, validation, repository-based discovery, and export to NCBI GEO submission format. FRED is designed for non-computational scientists and specialized facilities alike, and integrates into existing RDM infrastructure without requiring dedicated IT resources.

Jasmin Walter, C. Kuenne, Noah Knoppik et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.