Skip to content
Review

From Overload to Insights: How AI Agents Can Support Scientists in Analyzing Complex Data

Jul 2026 · arXiv.org · Vol abs/2607.16845 · 0 citations · 64 references
Computer Science

TL;DR

This study identifies key knowledge challenges in scientific data analysis, derives requirements for an AI agent that supports knowledge retrieval and source code generation, and proposes design recommendations for a specialized system adaptable to the evolving AI tool landscape.

Abstract

Scientists at European XFEL conduct experiments that generate very large and complex datasets. The subsequent data analysis is challenging as scientists must combine their domain expertise with facility- and software-specific knowledge scattered across documentation, tools, and support channels. To address this problem, we designed and evaluated an agentic AI system tailored to the scientists'needs and integrated with the high-performance computing environment of European XFEL. Using a design science research approach, we conducted a rapid literature review, a systematic evaluation of 16 AI tools, multiple interviews, a focus group, and a user study with experts at European XFEL to develop and evaluate two prototypes. Our study identifies key knowledge challenges in scientific data analysis, derives requirements for an AI agent that supports knowledge retrieval and source code generation, and proposes design recommendations for a specialized system adaptable to the evolving AI tool landscape. These findings provide guidance for developing maintainable AI support in highly specialized scientific environments.

View source

Similar papers

Review Open access 2024

Agentic AI Framework for Autonomous Scientific Research Assistance

Experimental results indicate that coordinated autonomous agents significantly reduce research time, improve workflow consistency, enhance knowledge discovery, and increase scientific productivity compared with conventional AI-based research assistants.

Anatoly Kitov, M. Kartsev · 0 citations
Jul 2026

An Interpretable AI Architecture for System-Level Reasoning from Engineering Documentation

Engineering systems are increasingly characterized by large, heterogeneous collections of technical documentation, including specifications, interface descriptions, and contribution records. While artificial intelligence techniques have been applied to document analysis, many existing approaches rely on opaque models that limit transparency and human trust. This paper presents a structured AI-based approach for deriving systemlevel understanding from engineering documentation by combining semantic abstraction, modular reasoning, confidence-aware outputs, and analyst validation. The approach emphasizes transparency and evidence-linked reasoning, enabling users to inspect intermediate representations and validate inferred relationships. A case study using a large-scale wireless systems documentation corpus and a focused Wi-Fi Aware worked example demonstrates how source-anchored reasoning can scale across extensive document sets while preserving human oversight.

Amrutha Moorthy · 0 citations
Review Open access Aug 2026

A New Paradigm: Agentic AI for Scientific Discovery

This article examines the emerging paradigm of agentic AI for scientific discovery, traces the conceptual shift from tools to agents, lays out a six-stage workflow spanning literature synthesis to manuscript generation, and reviews practical systems in chemistry, equation discovery, materials science, and general machine learning research.

Alexander Taktakidze · 0 citations
Book Aug 2026

KDD AI reasoning day

Large language models and foundation models are increasingly embedded in reasoning systems that plan, invoke tools, use memory, gather evidence, and iteratively refine their outputs. The second KDD Day on AI Reasoning brings together researchers and practitioners from academia and industry to examine how these systems can be made more capable, reliable, interpretable, and efficient. The program spans scientific discovery, human-centered interaction, software engineering, time-series analysis, deep research, computer use, and inference infrastructure. Across these domains, the day highlights shared challenges: grounding decisions in evidence, designing effective feedback and verification mechanisms, evaluating open-ended behavior, managing test-time computation, and preserving meaningful human control. Through keynote and invited presentations, the event provides a forum for connecting advances in models, agents, data, systems, and applications, and for identifying research directions toward trustworthy next-generation reasoning systems.

Jun Huan, James Caverlee, Lei Li et al. · 0 citations
Review Aug 2026

MUSE: An Interactive Meta-Agent for Understanding and Steering LLM-powered Data Science Systems

MUSE is presented, an interactive meta-agent that enhances user understanding and control of agentic data science systems by dynamically restructuring low-level execution traces into multiple semantic levels that support navigation from high-level overviews to low-level implementation details.

Wei-Hao Chen, Weixi Tong, Yuan Tian et al. · 0 citations
Preprint Aug 2026

AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups

Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific workflows. However, using proprietary commercial AI systems raises concerns about transparency, reproducibility and privacy, which are essential for scientific practices. To this end, AquiLLM was developed as an open-source modular RAG-LLM framework using open-weight models, designed to support research groups in capturing tacit knowledge. In this work, we present a series of architectural improvements and feature enhancements to AquiLLM, including local embedding and reranking, multimodal capabilities, OpenAI-compatible inference interfaces, user interface improvements, semantic and episodic memory capabilities, and skills support. These enhancements were informed by discussions with domain experts, including astrophysicists and environmental researchers, and represent a step toward AI systems more closely aligned with scientific research practices.

J. Stark, S. Saikrishnan, Vikram Seenivasan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.