Skip to content

Agentic AI for Gravitational Wave Data Analysis: A Head-to-Head Comparison of Coding Agents Executing a Matched Filter Pipeline on Einstein Telescope Simulated Data

Gianluca Inguglia
Sep 2026
Artificial Intelligence Human-computer Interaction

Abstract

We report a methodological study of agentic AI in gravitational-wave data analysis: two systems, Claude Code (Anthropic) and Codex (OpenAI), autonomously executed the same simple end-to-end pipeline on Einstein Telescope (ET) simulated data, on shared infrastructure and without human intervention. The object of study is the behaviour, reliability and auditability of the agents, not the physics output, used here as a controlled test case. The pipeline comprises power spectral density estimation from simulated ET noise, geometric template bank generation with IMRPhenomD waveforms, matched-filter recovery of 100 binary black hole injections, results generation, and LLM-assisted production of a LaTeX manuscript in Physical Review D style. Both agents received identical specifications and resources. The experiment was run twice: first with unrealistically loud injections, then with signals rescaled to a physically motivated SNR range. In both runs the results converged, with comparable detection efficiency and template bank size. The agents, however, behaved very differently: Claude Code finished in about 3.4 minutes with silent deviations from the specification, while Codex needed about 16 minutes across explicit self-correcting restarts, including an unsolicited optimization of the matched-filter inner loop. In the second run, a subtle difference in interpreting the SNR-range instruction produced a genuine scientific divergence: Claude Code silently raised the SNR floor to 8 (100% efficiency), while Codex followed the specification literally down to SNR 7 and recorded one missed detection. We discuss the implications - speed versus auditability, silent deviation versus explicit self-correction, instruction interpretation, and intermediate data representations in multi-model pipelines - for agentic AI in scientific workflows, within the limits of a single-pipeline, two-run benchmark.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.