This study identifies key knowledge challenges in scientific data analysis, derives requirements for an AI agent that supports knowledge retrieval and source code generation, and proposes design recommendations for a specialized system adaptable to the evolving AI tool landscape.
Abstract
Scientists at European XFEL conduct experiments that generate very large and complex datasets. The subsequent data analysis is challenging as scientists must combine their domain expertise with facility- and software-specific knowledge scattered across documentation, tools, and support channels. To address this problem, we designed and evaluated an agentic AI system tailored to the scientists'needs and integrated with the high-performance computing environment of European XFEL. Using a design science research approach, we conducted a rapid literature review, a systematic evaluation of 16 AI tools, multiple interviews, a focus group, and a user study with experts at European XFEL to develop and evaluate two prototypes. Our study identifies key knowledge challenges in scientific data analysis, derives requirements for an AI agent that supports knowledge retrieval and source code generation, and proposes design recommendations for a specialized system adaptable to the evolving AI tool landscape. These findings provide guidance for developing maintainable AI support in highly specialized scientific environments.
Experimental results indicate that coordinated autonomous agents significantly reduce research time, improve workflow consistency, enhance knowledge discovery, and increase scientific productivity compared with conventional AI-based research assistants.
Anatoly Kitov, M. Kartsev· International Journal of Eme...· 0 citations
Engineering systems are increasingly characterized by large, heterogeneous collections of technical documentation, including specifications, interface descriptions, and contribution records. While artificial intelligence techniques have been applied to document analysis, many existing approaches rely on opaque models that limit transparency and human trust. This paper presents a structured AI-based approach for deriving systemlevel understanding from engineering documentation by combining semantic abstraction, modular reasoning, confidence-aware outputs, and analyst validation. The approach emphasizes transparency and evidence-linked reasoning, enabling users to inspect intermediate representations and validate inferred relationships. A case study using a large-scale wireless systems documentation corpus and a focused Wi-Fi Aware worked example demonstrates how source-anchored reasoning can scale across extensive document sets while preserving human oversight.
Amrutha Moorthy· International Symposium on C...· 0 citations
This article examines the emerging paradigm of agentic AI for scientific discovery, traces the conceptual shift from tools to agents, lays out a six-stage workflow spanning literature synthesis to manuscript generation, and reviews practical systems in chemistry, equation discovery, materials science, and general machine learning research.
Alexander Taktakidze· Longevity Horizon· 0 citations
Large language models and foundation models are increasingly embedded in reasoning systems that plan, invoke tools, use memory, gather evidence, and iteratively refine their outputs. The second KDD Day on AI Reasoning brings together researchers and practitioners from academia and industry to examine how these systems can be made more capable, reliable, interpretable, and efficient. The program spans scientific discovery, human-centered interaction, software engineering, time-series analysis, deep research, computer use, and inference infrastructure. Across these domains, the day highlights shared challenges: grounding decisions in evidence, designing effective feedback and verification mechanisms, evaluating open-ended behavior, managing test-time computation, and preserving meaningful human control. Through keynote and invited presentations, the event provides a forum for connecting advances in models, agents, data, systems, and applications, and for identifying research directions toward trustworthy next-generation reasoning systems.
Jun Huan, James Caverlee, Lei Li et al.· Proceedings of the 32nd ACM...· 0 citations
MUSE is presented, an interactive meta-agent that enhances user understanding and control of agentic data science systems by dynamically restructuring low-level execution traces into multiple semantic levels that support navigation from high-level overviews to low-level implementation details.
Wei-Hao Chen, Weixi Tong, Yuan Tian et al.· 0 citations
Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific workflows. However, using proprietary commercial AI systems raises concerns about transparency, reproducibility and privacy, which are essential for scientific practices. To this end, AquiLLM was developed as an open-source modular RAG-LLM framework using open-weight models, designed to support research groups in capturing tacit knowledge. In this work, we present a series of architectural improvements and feature enhancements to AquiLLM, including local embedding and reranking, multimodal capabilities, OpenAI-compatible inference interfaces, user interface improvements, semantic and episodic memory capabilities, and skills support. These enhancements were informed by discussions with domain experts, including astrophysicists and environmental researchers, and represent a step toward AI systems more closely aligned with scientific research practices.
J. Stark, S. Saikrishnan, Vikram Seenivasan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.