Every memory-based knowledge editor in the SERAC lineage depends on a scope decision: given a query, does a stored edit apply? We report that current knowledge-editing benchmarks cannot measure this decision at all. Using INLAY, a gradient-free editor we built to obtain exact per-query ground truth (the model is frozen, edits live in an external addressable memory, and applying an edit is a bias added along one token's unembedding direction at decode time), we execute every candidate router action on 1,689 queries spanning three datasets and three input conditions. An oracle router choosing the best action every time ties a one-line static policy to four decimal places in all nine dataset-by-condition cells: the maximum attainable gain of any per-query routing method is 0.00 points. Abstention is the sole winning action zero times out of 1,689. The cause is structural: these are counterfactual benchmarks whose evaluation question asks for the post-edit answer, so answering from parametric knowledge is wrong by construction, and a benchmark without negatives cannot reward a classifier's ability to reject. This generalizes beyond our system to the whole scope-classifier family the benchmarks are used to evaluate. We confirm the mechanism directly: constructing the missing condition ourselves, by withholding a query's own edit from the index for half the sample, moves pooled headroom from exactly +0.0000 to +0.0420 and gives abstention its first wins. We also report where INLAY itself does not win (WISE beats it on Qwen2.5-7B CounterFact, and retrieval-augmented generation beats every method we tested, INLAY included, on rigorously matched RippleEdits), and disclose two bugs found during a self-audit of our own routing machinery, neither of which changed a published headline number outside noise.
Background Generative Artificial Intelligence (GAI) models, such as GPT-4, have been extensively studied for their integration into medical practice and education. GPT-4 has demonstrated excellent performance on medical licensing examinations, including the United Kingdom Medical Licensing Assessment (UKMLA). However, the field lacks longitudinal data on whether such performance is stable or varies over time. Theoretically, iterative improvements updated by the vendor should enhance performance, but empirical evidence of such longitudinal trends remains limited. Given that GPT-4 undergoes periodic vendor-side updates, we aimed to analyse the categorical, time-spaced performance of GPT-4 on the UKMLA to characterise how its performance changes over time in a medical context. Methods Two publicly available UKMLA papers were fed into GPT-4 at two different time points, June 2023 and December 2024. 191 questions were provided with and without multiple-choice options to assess GPT-4’s clinical competence. McNemar’s test was performed to evaluate changes in GPT-4’s performance over time, comparing domain-specific questions. Results GPT-4’s accuracy improved noticeably between the two rounds (MCQ: 88.0 to 93.7%, p = 0.027; non-MCQ: 68.1 to 81.7%, p = 1.00). Single-step accuracy rose from 73.1 to 82.3%, and multi-step from 57.4 to 80.3% without MCQ. GPT-4 showed improved accuracy from Round 1 to Round 2 for both single-step and multi-step questions, with MCQ-prompted responses consistently outperforming non-MCQ responses (up to 95.1% accuracy for multi-step MCQ questions in Round 2). GPT-4’s performance improved across all question categories from round one to round two, most notably in management questions without MCQ options (+23.30%), though these differences were not statistically significant. Discussion and conclusion GPT-4’s performance on the UKMLA improved significantly over 18 months, suggesting that iterative vendor-side model updates enhance clinical reasoning capabilities. These findings indicate that GPT-4 may serve as a supplementary educational tool for medical students and clinicians; however, the underlying drivers of performance changes remain opaque, and such tools should be deployed with structured oversight to prevent overreliance.
Ravanth Baskaran, Sai Sirikonda, Aditya Singh et al.· Frontiers in Medicine· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.