This paper introduces a novel multi-AI-agent system designed to fully automate SLRs, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision.
Abstract
Systematic literature review (SLR) is foundational to evidence-based research, enabling scholars to identify, classify, and synthesize existing studies to address specific research questions. Conducting an SLR is, however, largely a manual process. In recent years, researchers have made significant progress in automating portions of the SLR pipeline to reduce the effort and time required for high-quality reviews; nevertheless, there remains a lack of AI-agent-based systems that automate the entire SLR workflow. To this end, we introduce a novel multi-AI-agent system designed to fully automate SLRs. Leveraging large language models (LLMs), our system streamlines the review process to enhance efficiency and accuracy. Through a user-friendly interface, researchers specify a topic; the system then generates a search string to retrieve relevant academic papers. Next, an inclusion/exclusion filtering step is applied to titles relevant to the research area. The system subsequently summarizes paper abstracts and retains only those directly related to the field of study. In the final phase, it conducts a thorough analysis of the selected papers with respect to predefined research questions. This paper presents the system, describes its operational framework, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision. The code for this project is available at: https://github.com/GPT-Laboratory/SLR-automation .
Systematic reviews, scoping reviews, mapping studies, and related evidence syntheses are increasingly difficult to conduct with fully manual workflows as search volumes, update cycles, and synthesis requirements continue to expand. At the same time, artificial intelligence, machine learning, and large language models are rapidly entering review practice across query formulation, screening, extraction, categorization, appraisal support, and reporting. Yet the empirical evidence remains uneven, task-dependent, and insufficient to justify unconstrained automation. Existing standards such as PRISMA 2020, PRISMA-S, PRISMA-ScR, PRISMA-P, PRESS, and SWiM remain essential, but none provides an end-to-end operational standard for when AI use is methodologically appropriate, how it should be validated, which review decisions must remain human-led, and how AI involvement should be reported so that readers can audit it. This paper proposes ARISMA, an AI Reporting and Integration standard for Systematic Methods and Analysis. ARISMA treats AI as an inspected, benchmarked, logged, and reversible assistant rather than an autonomous reviewer. It is built around one governing principle: every consequential scientific decision must remain human-interpretable, human-auditable, and human-accountable. The paper contributes a lifecycle taxonomy, process guidance, stepwise recommendations across the review pipeline, a governance and provenance model, a tool-support framework, an AI-integrated reporting checklist, and a validation matrix. It also addresses legal, privacy, infrastructure, and sustainability considerations. The framework was iteratively refined through structured expert consultation. The result is a practical and auditable guideline for responsible AI-assisted evidence synthesis.
Systematic literature review of clinical trials drives regulatory decision-making, but conventional screening and extraction are time-consuming, labor-intensive, and vulnerable to study selection bias. We propose two fit-to-purpose multi-agentic systems (MAS) for systematic literature review, with human-in-the-loop. The screening MAS uses multiple LLM agents with heterogeneous personas and multiround cross-review, and uniformly improves accuracy over a single-LLM baseline. The extraction MAS combines standardization, an iterative correction loop, and retrieval-based context control to ensure accuracy and scalability. Both MAS are specifically designed to support Human-In-The-Loop which is essential for clinical decisions. The novelty of the proposed approach lies in the system architecture rather than in any single foundation tools: the system can naturally benefit from future improvements in the underlying tools, for instance, stronger LLM agents, retrieval engines, image recognition methods, etc. As a real-world application, a published network meta-analysis is reproduced by the MAS. The result recovers all trials from the original study and identifies additional eligible trials missed by manual review, leading to updated clinical conclusions.
Zexin Ren, Zixuan Zhao, Qiyun Li et al.· 0 citations
AI tools increasingly support tasks across the scientific research cycle, from experiment design and manuscript preparation to peer review. At the same time, the continuing growth in conference submissions has increased the burden on meta-reviewers, who must synthesize reviewer feedback, author rebuttals, and manuscript revisions. To address this concern, this paper introduces Metag, a dataset to accelerate the development of meta-reviewing agents, specifically to identify changes made to scientific articles during the review-rebuttal process. Each instance contains a reviewer concern, the author's proposed resolution, and the manuscript diffs implementing the stated change. Metag is collected by obtaining manuscript versions from before the review deadline and after acceptance, computing differences between the two documents, and asking human annotators to align these differences with action items from OpenReview discussions. The resulting dataset consists of 349 high-quality action items tied to paper differences and will enable building methods to empower meta reviewers to quickly identify whether authors have addressed reviewer statements and where in the paper those changes have been made, resulting in additional transparency and traceability throughout peer review. The dataset is publicly available at https://github.com/microsoft/Metag-dataset.
Anirudh S. Sundar, Min Chen, Divya Tadimeti et al.· 0 citations
Automatic scientific survey generation has become an important task in scientific document processing. The common approach of retrieving literature from a single source (e.g., arXiv) and generating surveys through a one-pass large language model (LLM) call often leads to limited reference coverage and, more importantly, fails to replicate the expert-driven revision process that is crucial for writing high-quality surveys. In this paper, we introduce SurveyAgent-HKA, a multi-agent framework that improves end-to-end scientific survey generation by incorporating knowledge derived from published surveys and peer-review comments. The framework decomposes survey generation into well-defined sub-tasks handled by LLM-powered agent. It first retrieves relevant papers from multiple sources and identifies key topics through clustering to construct an initial outline, which is then refined using outlines from related human-written surveys. Based on the refined outline, topic-focused papers are retrieved and re-ranked to select for drafting a well-grounded survey. Then, we identify common issues raised by experts in peer-review comments from published surveys to guide the revisions and finalize the survey. Experiments on two domains show that our approach outperforms mainstream baselines in citation quality, structural consistency, and content quality. Furthermore, our framework is efficient in both time and cost, making it a practical solution for broader AI-assisted scientific writing applications.
Tong Bao, Mir Tafseer Nayeem, Yi Zhao et al.· 0 citations
Inspired by search and recommender systems, this work builds Find, Attempt, and Recommend (FAR), a literature-to-review cascade that automates the search for suitable problems and focuses human attention on artifacts that have passed several stages of filtering.
Zeyu Zheng, Shengtong Zhang, Jeremy Avigad et al.· 0 citations
ABSTRACT Introduction Artificial intelligence (AI) is a branch of technology enabling machines to emulate complex human skills; it can also entail problem‐solving using bioinspired methods. It is used for automating systematic literature reviews (SLR), that is, defining a clinical question, locating relevant literature, preliminary screening, study evaluation, data extraction and analysis. Title and abstract screening is one of the most time‐consuming and error‐prone phases involved in developing a systematic review. While AI promises to expedite this process, adopting it faces challenges due to concerns about compatibility and transparency. This review aims to identify current evidence concerning AI use during preliminary SLR reference screening; it describes characteristics such as the different metrics used for reporting performance and how the different algorithms, pipelines, workflows or web applications are validated. AI resource users' reflections regarding SLR screening automation have also been summarized. Methods A scoping review was conducted following Joanna Briggs Institute's (JBI) methodology. Its objective was to identify existing evidence regarding the use of AI resources for title and abstract screening automation. Searches were limited to articles published between 2019 and 2026. The review included primary studies reporting the development, assessment, validation, or real‐world use of AI resources for screening automation, as well as systematic reviews and articles reporting experiences or recommendations for their use. Two types of data were extracted: (1) from primary studies—characteristics of AI resources and, where applicable, recommendations for their use; (2) from systematic reviews and experience‐based articles, recommendations for the use of AI resources. Results included frequency descriptions, tables, figures, and a decision flowchart reflecting the number of references and articles retrieved, excluded, or included in the final analysis. Results A total of 174 unique studies published between 2019 and 2026 were included in this scoping review. These were grouped into web applications (43%), model comparisons (32%), generative models (6%), pre‐trained models (3%) or pipelines/workflows (15%) used for title and abstract screening in systematic literature review (SLR). Most studies came from North America. Evaluating these tools often relied on retrospective comparisons with human reviewers' work (63%), sensitivity (n = 60), and specificity (n = 62) being the most reported metric for criterion assessment and Work Saved over Sampling (n = 28) being the most reported metric for assessing their utility. Considerations concerning AI resource use focused on the need for standardized evaluation metrics, stopping criteria, study design and the data sets used, resource characteristics facilitating usability, best practice and future research areas, with the persistence of the human component in the process (n = 26) being the most pressing recommendation. Conclusion The findings indicated substantial heterogeneity regarding the types of AI resources used, considerable variation concerning the metrics used for reporting performance, differences in how such metrics are defined and a clear need for standardizing reporting methods, study designs and related procedures. Although AI technologies will continue to evolve, maintaining a clear and consistent framework for interpreting research on AI resources for automating title and abstract screening can support understanding their level of maturity and facilitate informed decision‐making by users.
A. M. Barragán, Sara Elena Ortiz Bonett, Eliana-Isabel Rodríguez-Grande et al.· Cochrane evidence synthesis...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduJun 3, 2026
The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.
MIT News · Artificial Intelligence· news.mit.eduAug 18, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.
MIT News · Artificial Intelligence· news.mit.eduJul 14, 2026
Through research and entrepreneurship, Professor Devavrat Shah is helping to design methods that can handle constant decision-making using limited computational resources.
Short chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, human-editable specifications. The post Flint: A visualization language for the AI era appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.