Artifacts of the paper "Tool Recommendation Using LLMs for Scientific Workflow Development: A Case Study of Galaxy and Nextflow"
Abstract
This replication package accompanies the paper “Tool Recommendation Using LLMs for Scientific Workflow Development: A Case Study of Galaxy and Nextflow”. The study investigates the capability of general-purpose Large Language Models (LLMs); GPT-5.5, Gemini 3.6 Flash, and DeepSeek-V3 to recommend platform-specific components for scientific workflow development in Galaxy and Nextflow/nf-core. The evaluation covers ten representative scientific workflow tasks, with five tasks for each workflow system, and assesses recommendation validity, essential-operation completeness, platform-related errors, and workflow-level compatibility. The quantitative evaluation is complemented by feedback from experienced scientific workflow developers. A core component of this package is the expert-curated reference sets used to evaluate LLM-generated recommendations. Because a scientific workflow task may have multiple valid implementations, the study does not assume a single workflow or tool sequence as the only correct solution. For each task, the reference sets identify the essential analytical operations, corresponding Galaxy tools or canonical nf-core modules, functionally valid alternatives, optional components, and valid analytical routes. The sets were developed using documented workflows and official Galaxy and nf-core documentation and were further reviewed and refined through expert assessment. Recommendations were categorized as exact matches, functionally valid alternatives, useful optional components, unnecessary or irrelevant components, platform-incompatible components, or unsupported/nonexistent components. The package also includes the Galaxy and Nextflow survey instruments used for the experienced-developer evaluation. The surveys contain participant information and consent material, demographic and experience questions, prompting guidance, the structured prompt template, workflow tasks, and questions for assessing LLM-generated tool or module recommendations. The study received approval from the Research Ethics Board of the anonymized institution, participants provided informed consent, and responses were handled anonymously and reported in aggregate. To preserve anonymity during peer review, identifying information such as the university name, researcher and supervisor names, institutional details, email addresses, and survey contact information has been removed or replaced with anonymous placeholders. The methodological content required to understand and replicate the study has been retained.This replication package is intended to support the reproduction, verification, and extension of the study. Researchers can use the provided workflow tasks and prompting materials to reproduce LLM recommendations, the expert-curated reference sets to assess recommendation validity and completeness, and the anonymized survey instruments to understand the experienced-developer evaluation procedure. Because LLMs and scientific workflow ecosystems evolve over time, exact recommendations may differ when the study is repeated with later model versions or updated Galaxy and nf-core components.