PROVE (Programmatic Reporting and Output Verification Engine), an auditable framework that uses optional LLM and retrieval support for table interpretation while reserving numerical and logical decisions for programmed validators, is introduced.
Abstract
Ensuring the accuracy and consistency of clinical trial Tables, Figures, and Listings (TFLs) remains a major challenge in regulatory reporting. Independent programming and manual review are essential quality-control practices, but cross-output verification still depends heavily on reviewer inspection and may miss structural, logical, or arithmetic discrepancies. Large language models (LLMs) can help interpret varied table language and navigate lengthy study documents, but they are not reliable substitutes for programmed statistical checks. We introduce PROVE (Programmatic Reporting and Output Verification Engine), an auditable framework that uses optional LLM and retrieval support for table interpretation while reserving numerical and logical decisions for programmed validators. PROVE links findings to source evidence, supports cross-output consistency checks, and allows LLM use to be enabled or disabled based on study requirements. We evaluated PROVE using ten replicated synthetic oncology reporting packages generated from raw data through SDTM, ADaM, and TFL outputs, with paired clean and discrepancy-injected packages; each replicate included 15 randomly injected discrepancies. We examined two table-label settings: exact labels matching the validator vocabulary and labels with similar clinical meaning but different wording. Within the implemented rule classes, all automated PROVE variants achieved perfect classification in the exact-label setting. In the label-variation setting, LLM-assisted semantic matching improved overall recall from 0.588 to 0.993 and overall F1 from 0.735 to 0.996 compared with exact-match, fuzzy lexical, and embedding-similarity variants. These findings suggest that LLMs are most useful for interpreting real-world variation in TFL wording and formatting, while executable checks should remain responsible for final numerical validation.
Background/Objectives: Obtaining timely access to detailed clinical trial data is not always straightforward. Privacy requirements, governance processes, and study-specific eCRF configurations can delay access, particularly during study start-up, when teams need realistic data to develop and test validation rules, repo...
Szymon Musik, Jacek Zalewski, Julia Jurkowska et al.· Healthcare· 0 citations
Evidence-based medicine demands strict logical consistency, yet current evaluations of large language models (LLMs) prioritize superficial label matching over genuine reasoning. We introduce LogiMed-RoB, a benchmark grounded in Cochrane Risk of Bias (RoB) 2.0 expert logic, comprising 860 randomized controlled trials (R...
Jia-Yu Huang, Zi-Chen Tang, Qian-Hui Ling et al.· 0 citations
Background: Large language models (LLMs) can support systematic reviews, but accurate individual outputs do not establish whether the final synthesis preserves the clinical question, accounts for statistical dependence and incorporates corrections. Objective: To develop and evaluate a framework linking LLM-assisted evi...
Statistical analysis of clinical data requires expertise in medical statistics. Large language models (LLMs) are increasingly used for code generation and may support both descriptive and advanced analyses, but their reliability remains uncertain. This study evaluated whether five current LLMs (GPT 5.3, Claude Sonnet 4...
J. Sam, T. Spreuer, M. Berger et al.· Studies in Health Technology...· 0 citations
Surrogate endpoints are widely used in clinical trials to accelerate treatment evaluation, yet their validity may vary substantially across patient subgroups. Although recent advances in heterogeneous causal mediation analysis enable subgroup-specific surrogate evaluation, applying these methods requires substantial ex...
Systematic reviews underpin clinical guidelines, yet their data-extraction step is a major expert-labor bottleneck bound by a protocolized workflow: two reviewers extract each study independently, an adjudicator resolves disagreements, and the team keeps an auditable record of how every value was produced. Large langua...
S. Kosuri, A. Bhosale, M. Glick et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.