Skip to content
Book Open access

The 1st International Workshop on AI Data Scientist

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 13423-13424 · 0 citations · 16 references

Abstract

As data volumes and analytical demands grow, traditional data science workflows struggle to meet the need for efficiency, scalability, and reliability. The rapid advancement of large language models (LLMs) has opened new possibilities for AI-powered agents to augment or automate end-to-end data science pipelines—from data exploration and cleaning to modeling, evaluation, and deployment. This emerging paradigm, termed the AI Data Scientist, has gained significant attention in research and industry, yet discussions remain fragmented regarding its integration, evaluation, and real-world impact. This workshop seeks to consolidate these efforts by providing an interdisciplinary forum for presenting cutting-edge research, sharing deployment experiences, and showcasing real-world systems. The workshop will feature invited talks, paper presentations, a demo track, and a panel discussion, aiming to foster community-building and guide responsible development in this rapidly evolving field.

Read PDF

Similar papers

Preprint Jul 2026

AI-Ready Research Workflows in Computational Social Science: Lessons on Building a Shared Language for Interdisciplinary Collaboration

Artificial intelligence (AI) is gaining traction in the social sciences and humanities (SSH). However, adoption remains limited by technical barriers to high-performance computing (HPC), validation processes that lag behind AI's rapid progress, and reproducibility standards that most SSH teams cannot meet. Research workflows--common in the life sciences--address these problems via encoding and abstracting technical complexity into repeatable routines; yet, accounts of how to build them in SSH remain scarce. We report on a two-year effort to build a workflow that enables a Science and Technology Studies unit to query, analyze, and enrich OpenAlex--a database of some 460 million scholarly records--on the MareNostrum supercomputer, using methods ranging from large-scale bibliometrics to LLM-based classification. We found the main challenge was translating domain-specific research questions into engineering requirements -- bridging two distinct methodological languages, with implications that were both organizational and technical. Organizationally, it meant adopting and adapting Agile to the research rhythm and pace, and reframing collaboration from a service arrangement to a co-design process. Technically, model-driven engineering was as valuable for collaboration as it was for automation; co-building the model facilitated both the creation of a shared vocabulary and the abstraction of HPC complexity. Finally, we highlight limitations we found in validation, reproducibility, and FAIR metadata -- beyond what any single project can sustain -- calling for coordinated, cross-institutional investment in the tooling and standards needed for AI-ready SSH workflows sustainable at scale.

J. Giner-Miguelez, A. Malaga, Felipe L Gómez-Cortés et al. · 0 citations
Preprint Aug 2026

AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups

Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific workflows. However, using proprietary commercial AI systems raises concerns about transparency, reproducibility and privacy, which are essential for scientific practices. To this end, AquiLLM was developed as an open-source modular RAG-LLM framework using open-weight models, designed to support research groups in capturing tacit knowledge. In this work, we present a series of architectural improvements and feature enhancements to AquiLLM, including local embedding and reranking, multimodal capabilities, OpenAI-compatible inference interfaces, user interface improvements, semantic and episodic memory capabilities, and skills support. These enhancements were informed by discussions with domain experts, including astrophysicists and environmental researchers, and represent a step toward AI systems more closely aligned with scientific research practices.

J. Stark, S. Saikrishnan, Vikram Seenivasan et al. · 0 citations
Book Open access Aug 2026

Automating End-to-End Hybrid Query Processing: Benchmark, Solution, and Insights

Hybrid queries—natural language questions over structured data that require both database capabilities and LLM reasoning—have recently emerged as a prominent research topic. However, existing solutions remain overly dependent on manual workflows, and current benchmarks are limited in scale and diversity. To bridge this gap, we present (1) HyQBench \xspace, a large-scale benchmark with 60\sim 90× more queries than prior work, built on 3× more databases; (2) AutoHyQ \xspace, an automated pipeline that can execute existing methods without manual intervention; (3) multi-dimensional, fine-grained evaluation metrics for comprehensive assessment. Through extensive experiments across multiple hybrid query approaches on diverse LLM backbones, we reveal their strengths and limitations, and identify research opportunities for advancing this emerging field. Our code and data are available at https://github.com/XMUDM/HyQBench.

Bo Li, Chenzhan Wang, Long-Kang Lin et al. · 0 citations
#generative ai Review Open access Sep 2026

Opportunities and challenges of generative AI in the research lifecycle

Artificial intelligence (AI) is increasingly being explored and adopted across the research lifecycle, from idea generation and literature discovery to data analysis, manuscript preparation, and editorial and peer-review processes. This perspective provides an overview of AI’s role across the stages of scientific research. We describe emerging tools and workflows, illustrating how AI can assist researchers by aggregating and synthesizing a large body of work across various domains, supporting methodological implementation, and facilitating communication and publication. We also discuss their shortcomings, including surface-level reasoning, the fabrication of plausible but incorrect outputs, and the challenges posed by the fact that researchers new to a field may not know which questions to ask or which nuances to interrogate. In addition, we discuss recent advances toward more agentic and end-to-end AI systems, highlighting both their technical feasibility and the challenges they pose for validation, oversight, and responsible use. For each stage of the research lifecycle, we outline key limitations of current AI systems and propose practical considerations, what researchers should and should not do to support rigorous and ethical integration of AI into scientific workflows. This integration requires coordinated frameworks across the ecosystem. Journals, funding agencies, universities, and policymakers play essential roles in defining standards for transparency and accountability, while individual researchers remain responsible for methodological rigor and validity of reported results.

Neda Sadeghi, Erin Nakamura, Luke J. Norman et al. · 0 citations
Open access 2018

Building Scalable Data Infrastructure for Generative AI Models: Challenges and Solutions

The rapid advancement of Generative AI models has underscored the necessity for robust and scalable data infrastructures capable of managing vast datasets and complex computational requirements. This paper explores the unique challenges encountered in building such infrastructures, including data acquisition, storage, processing, and real-time access. We analyze existing solutions and propose best practices for designing architectures that ensure efficiency, scalability, and reliability. By examining case studies and current industry practices, the paper provides a comprehensive framework for developing data infrastructures tailored to the demands of Generative AI applications.

Kwame Nkosi · 0 citations

Improving Data Preparation for CSV Files with LLMs

This work shows how machine learning and GenAI can be used to assist with two specific tasks: First, when reading CSV files, it needs to be decided whether the first row is a header or not, and how machine learning and GenAI can be used to assist with two specific tasks.

Alexander van Renen, Moritz Rengert, Macallyster Edmondson et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.