Skip to content
Open access

Scientific discovery in the age of AI and supercomputing

Nov 2025 · Scientific Reports · 1 citation
Computer Science

TL;DR

Drawing on metadata from more than five million scientific publications, this work examines how the convergence of AI and HPC correlates with scientific breakthroughs and shows that this computational synergy is most pronounced at the scientific frontier.

Abstract

Artificial intelligence (AI) and high-performance computing (HPC) are transforming scientific capabilities and the way science is conducted. Yet their combined impact on scientific discovery remains poorly understood, as do inequalities in access to these capabilities across countries and institutions. Drawing on metadata from more than five million scientific publications (2000–2024) across 27 fields, we examine how the convergence of AI and HPC correlates with scientific breakthroughs. Our results show that this computational synergy is most pronounced at the scientific frontier: research combining AI and HPC is more likely to introduce novel ideas and achieve top-cited status than either conventional work or research using AI or HPC in isolation. We also document growing disparities in access to supercomputing resources and AI expertise, which are increasingly concentrated in a small number of regions (dominated by the United States and China, though the EU27 aggregate maintains high competitiveness in combined AI+HPC output). The future of discovery will depend not only on advances in algorithms and computing power, but also on enacting policies that democratise these capabilities across the global scientific ecosystem.

Read PDF

Similar papers

Open access Aug 2026

Scientific computing in the age of agentic AI: an exploratory field report

Overall, it is found that the use of coding agents in scientific computing holds great promise for accelerating scientific research and increasing the reliability of critical systems, but that outstanding concerns remain.

Jeremiah H. Li, Alex Rubinsteyn, Sergey Feldman et al. · 0 citations
Jul 2026

Ascend to Science: Exploration of AI Chips for Scientific Computing

The rapid rise of AI-oriented accelerators has reshaped compute systems around low-precision tensor engines, raising a practical question for the HPC community: under what conditions can such hardware support scientific workloads that demand numerical robustness, irregular memory access, and scalability? Using the Ascend 910 NPU series as a representative tensor-centric platform, we characterize precision, execution, and memory-hierarchy bottlenecks that hinder the direct deployment of scientific codes. We then develop and evaluate workload-specific mappings across five application studies -- HPL-MxP, LRSVD, SGEMM-cube, PQSim, and SMC-X -- combining heterogeneous execution, mixed-precision numerical formulations, precision emulation, hierarchical memory orchestration, and communication--computation overlap. These studies show that AI-native NPUs can achieve numerical robustness, competitive performance, and satisfactory scalability when numerical formulation, execution placement, and data movement are addressed in a coordinated manner. Our results provide a state-of-the-practice case study of how scientific workloads can be adapted to tensor-centric architectures, while distinguishing transferable optimization principles from Ascend-specific implementation details.

Weicheng Xue, Kai Yang, Yongxiang Liu et al. · 0 citations
Review Open access Aug 2026

Quo Vadis? Scientific Discovery in the Age of Artificial Intelligence

The advancement of AI in science raises broader questions concerning the division of cognitive labour between human researchers and machines, and a typology of AI systems used in research is proposed, ranging from specialised scientific AI through scientific AI assistants and agents to hybrid experimental systems that combine computation and physical experimentation.

P. Jedlička · 0 citations
Preprint Jul 2026

AI-Ready Research Workflows in Computational Social Science: Lessons on Building a Shared Language for Interdisciplinary Collaboration

Artificial intelligence (AI) is gaining traction in the social sciences and humanities (SSH). However, adoption remains limited by technical barriers to high-performance computing (HPC), validation processes that lag behind AI's rapid progress, and reproducibility standards that most SSH teams cannot meet. Research workflows--common in the life sciences--address these problems via encoding and abstracting technical complexity into repeatable routines; yet, accounts of how to build them in SSH remain scarce. We report on a two-year effort to build a workflow that enables a Science and Technology Studies unit to query, analyze, and enrich OpenAlex--a database of some 460 million scholarly records--on the MareNostrum supercomputer, using methods ranging from large-scale bibliometrics to LLM-based classification. We found the main challenge was translating domain-specific research questions into engineering requirements -- bridging two distinct methodological languages, with implications that were both organizational and technical. Organizationally, it meant adopting and adapting Agile to the research rhythm and pace, and reframing collaboration from a service arrangement to a co-design process. Technically, model-driven engineering was as valuable for collaboration as it was for automation; co-building the model facilitated both the creation of a shared vocabulary and the abstraction of HPC complexity. Finally, we highlight limitations we found in validation, reproducibility, and FAIR metadata -- beyond what any single project can sustain -- calling for coordinated, cross-institutional investment in the tooling and standards needed for AI-ready SSH workflows sustainable at scale.

J. Giner-Miguelez, A. Malaga, Felipe L Gómez-Cortés et al. · 0 citations
Book Open access Jul 2026

A Program for Providing Accelerated Cybertraining for Emerging Scientists

The utilization of advanced cyberinfrastructure (CI) for scientific research is rapidly growing, catalyzed by the proliferation of artificial intelligence and machine learning (AI/ML)-enabled tools and research workflows. The NSF has facilitated this growth through programs like ACCESS and NAIRR, but many researchers are still overwhelmed when trying to conduct scientific workflows using advanced CI resources. To effectively utilize these resources, researchers must have a deep understanding of advanced computing technologies, the national CI ecosystem, data management practices, and discipline-specific scientific workflows. PACES (Providing Accelerated Cybertraining for Emerging Scientists) is a collaborative NSF-funded cybertraining program promoting the adoption of advanced CI by providing hands-on training to researchers seeking to incorporate these resources in their own scientific workflows. Since its inception in 2024, PACES has hosted two in-person workshops, hosted 78 virtual courses, and provided 18 asynchronous courses. The workshops and short courses cover a variety of topics, including effective utilization of high-performance computing (HPC) resources; programming in Python, R, and Julia; AI/ML; and domain-specific scientific workflows. This extensive range of topics spans researcher experience levels allowing PACES to provide training at any level, meeting researchers where they are to enable rapid adoption of advanced CI resources.

Wesley A. Brashear, Joshua Winchell, Halle Gray et al. · 0 citations
Preprint Aug 2026

BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks

BixBench3, a benchmark that measures the capacity of AI agents to process raw biological data through to scientific results, reveals that LLMs vary substantially in their ability to execute multiple sequential analysis steps coherently, manage large quantities of raw data, and work across scientific domains.

Zane Koch, A. Wassie, Javier Valdes-Aleman et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.