Skip to content
#data science #explainable ai Review Open access

Diagnostic accuracy of deep learning applied to liver MRI for hepatocellular carcinoma detection: a systematic review and HSROC meta-analysis of patient-level evidence

Oct 2026 · Egyptian Liver Journal · Vol 16 · 0 citations · 33 references
Hepatocellular Carcinoma Treatment and Prognosis

Abstract

Liver MRI is central to non-invasive hepatocellular carcinoma (HCC) diagnosis. Deep learning (DL) is increasingly studied for detection and interpretation, but the maturity, transportability, and patient-level diagnostic evidence for these systems remain uncertain. We searched MEDLINE, Embase, and Web of Science from 2010 through July 2025 (the exact original execution day was not retained) and updated the same strategies through 20 July 2026. Eligible studies evaluated stand-alone DL or radiologist-interpreted DL-assist MRI and provided reconstructable patient-level 2 × 2 data. Two reviewers independently selected studies, extracted data, and assessed QUADAS-2 and AI-reporting items. Four retrospective Chinese tertiary-centre cohorts (N = 698; three stand-alone systems and one radiologist-interpreted DL-assist system) were included; the update identified no additional eligible study. Model-derived exploratory pooled estimates were sensitivity 0.862 (95% CI 0.819–0.896) and specificity 0.916 (0.881–0.941); MCMC median LR + was 10.19 (95% interval 7.22–14.55), LR- was 0.151 (0.114–0.199), and diagnostic odds ratio was 67.59 (41.10–111.10). Between-study variance estimates lay at the boundary and the confidence and prediction regions were nearly identical; with four studies, this does not establish statistical or clinical homogeneity. Patient selection and flow/timing commonly had high or unclear risk of bias, and validation, threshold, blinding, explainability, workflow, and regulatory reporting were incomplete. DL applied to liver MRI yielded promising model-derived patient-level summaries, but four retrospective, geographically concentrated cohorts do not establish a transportable accuracy benchmark, radiologist augmentation, or clinical benefit. The principal finding is an evidence-maturity gap requiring prospective multi-centre reader studies, locked testing, transparent reporting, and evaluation of workflow and safety.

Read PDF

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Trajectory Balance: Improved Credit Assignment in GFlowNets

It is proved that any global minimizer of the trajectory balance objective can define a policy that samples exactly from the target distribution, and empirically demonstrate the benefits of the trajectories balance objective for GFlowNet convergence, diversity of generated samples, and robustness to long action sequenc...

Esmeralda S. Whitammer, Moksh Jain, Emmanuel Bengio et al. · 302 citations · ⚡60

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.