Generative artificial intelligence (AI) is rapidly transforming how scientific knowledge is produced, reviewed, and disseminated. In response, journals and publishing organizations have begun issuing policies to govern AI use in scholarly publishing. However, it remains unclear whether existing governance frameworks meaningfully address the risks AI introduces across the full publication pipeline. We conducted a narrative review of journal policies, publisher guidance, and recent analyses of AI governance in scientific publishing, complemented by direct examination of submission guidelines from high-impact, open-access, and regional medical journals. Our findings show that while most journals have converged on a narrow legal consensus (prohibiting AI authorship and requiring disclosure), yet governance remains fragmented and incomplete. Policies disproportionately target text generation by authors, while leaving critical domains under-regulated, including AI-assisted data analysis, peer review practices, enforcement mechanisms, and equity implications for researchers and reviewers globally. To synthesize these findings, we introduce the AI governance readiness levels, a five-level framework for assessing how well-equipped journals are to govern AI across the research and publication process. We further describe PRAIDE (Preparation, Representation, Attribution, Integrity checks, Dissemination, and Evaluation) as an illustrative architecture that integrates existing policies, integrity safeguards, and post-publication oversight into a coherent governance model. We argue that effective AI governance in scientific publishing cannot be achieved through static rules or journal-centric control alone. Instead, it requires a shift toward shared, adaptive oversight of the scientific publishing system, recognizing that responsibility for governing AI in science is collective, continuous, and inseparable from the public trust in research.
R. Abulibdeh, J. Arslan, S. Ordóñez et al.· MIT Science Policy Review· 1 citation
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore. Across seven cohorts on six public datasets spanning text (MedQA-USMLE, MedMCQA, MIMIC-CXR reports), imaging (NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert) and tabular ICU records (SUPPORT2), Gemini committees resist these cues in isolation (flip 5-16%), yet a socially plausible shortcut spreads: when two peers assert the same wrong answer, the holdout under test adopts it in 38% of cases, as does a false"pre-screen"system flag, on both capability tiers. Of three oversight agents, a gate cannot separate adoption from honest agreement (false-positive rate 100%); a same-lineage judge reading only the transcript flags adoption on text (precision 100%, recall 93%) but collapses onto the gate in imaging; a referee that privately re-queries the holdout transfers to imaging (77-88% precision, 13-21% false-positive rate). Tripling a cue's visual salience does not move contagion, whereas a second peer voice raises it by half again. Gaming a hidden rubric is near-silent: only 1/10 text and 1/134 imaging drifters name the rubric they moved toward. What games a committee is social plausibility, and only a referee independent of self-report catches it. Code: https://github.com/criticaldata/benchmaxxing
S. Ordóñez, Agastya Munnangi, Aldo Marzullo et al.· 0 citations
This narrative review outlines the limitations of traditional evaluation frameworks and proposes a paradigm shift toward adaptive, iterative, and context-specific assessment methodologies that can progress from a promising technology to reliable clinical tools that improve patient outcomes, support clinical decision-making, and uphold ethical standards in routine practice.
M. Fosset, Joris Pensier, Boris Jung et al.· PLOS Digital Health· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.