Skip to content

PREreview of "When Drift Detectors cry Wolf: False Alarm Rates in continuous ML Monitoring"

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)
Data Stream Mining Techniques

Abstract

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23128096. Summary This paper studies a problem that production ML teams feel but the literature mostly ignores: drift detectors are evaluated on detection accuracy under synthetic, single-shot shifts, while in production they run continuously across many features and time windows — where even small nominal false positive rates accumulate into frequent alerts and, eventually, alarm fatigue. The author simulates a 30-day continuous monitoring cycle on the Adult Income dataset (~30k samples, 14 features) and compares five widely used detectors — PSI, KS, MMD, LSDD, and adversarial validation — across batch sizes 50–500 and injected drift magnitudes (5–20 year shifts in the age feature), reporting false-positive days, true positive rate, and time to detection. Headline findings: PSI fires on ~30/30 days below ~200 samples per batch and stabilizes sharply above it; KS achieves the best stability–sensitivity balance (TPR up to 0.90 at the largest shift); MMD shows near-zero sensitivity under default settings; Bonferroni correction suppresses false alarms at the cost of sensitivity. Strengths The framing is the paper's core contribution. Reframing drift detection from a detection-accuracy problem to an operations-reliability problem is exactly the shift practitioners need. I audit ML designs for production readiness, and "the detector works" vs. "the detector is trusted by the on-call team" are different claims — this paper measures the second one. The alarm-fatigue motivation (Sculley et al.) is not decoration; the metric design follows from it. "False positive days out of 30" is a metric an SRE can reason about; it maps directly to pager load. This is a better evaluation currency for monitoring work than AUC-style summaries, and more papers in this space should adopt it. The practical-guidelines table (Table 1) is genuinely useful. The per-detector "practical implication" column — e.g., PSI only above ~200 samples/batch, KS as the reliable default for tabular monitoring — is the kind of artifact that survives the paper and ends up in runbooks. Too many monitoring papers stop at curves; this one tells you what to configure. Experimental transparency is good. Full appendices with detector configurations (binning, thresholds, kernel settings, permutation counts), the monitoring protocol, and complete result tables with mean ± std across 5 seeds. The PSI batch-size table (30.00±0.00 at batch 50 → 0.00±0.00 at batch 500) is striking precisely because the numbers are all visible. Major issues 1. Single dataset, single drift type — generality is the main limitation. Everything rests on Adult Income: tabular, 14 features, and drift injected into one feature (age) as a gradual univariate shift. Real production drift is frequently multivariate, correlated across features, sudden (upstream schema or pipeline changes), or semantic (label drift with stable inputs). The conclusion acknowledges this as future work, but the practical guidelines in Table 1 are written as deployment advice — they should carry an explicit scope warning that the ~200-sample PSI threshold and the KS-default recommendation are calibrated to one tabular dataset, not established as general rules. 2. Label-encoding categoricals before KS/MMD distorts the comparison. Appendix A states categoricals were label-encoded. A two-sample KS test on label-encoded categoricals is not meaningful — the encoding imposes an arbitrary ordinal structure the test then treats as real. MMD with a Gaussian kernel on label-encoded categories has the same problem. This choice may flatter detectors that are insensitive to the distortion and penalize ones that react to it. At minimum, the paper should report a sensitivity analysis with proper categorical handling (e.g., one-hot + appropriate tests, or feature-type-aware detector variants) so readers know whether the KS-vs-MMD ranking survives it. 3. No compute-cost analysis, and it is first-order for the recommendation. MMD and LSDD are run with 2000-sample reference subsets and 100 permutations; adversarial validation trains a classifier per check. At 14 features × daily runs this is already non-trivial, and production tables routinely have hundreds of features. If KS wins on the stability–sensitivity trade-off and is orders of magnitude cheaper per check, that strengthens the recommendation considerably; if the kernel methods' cost is the real reason teams avoid them, the paper should say so. Wall-clock time per detector per batch size belongs in Appendix B. 4. Bonferroni is tested but FDR is not, despite being cited. The paper cites Benjamini–Hochberg (1995) in the related work, then evaluates only the Bonferroni correction — the most conservative multiple-testing adjustment available — and concludes corrections cost sensitivity. An FDR-controlling procedure is the natural middle ground for a monitoring setting where some false alarms are tolerable but alarm floods are not, and it is what many production teams would actually reach for. The stability–sensitivity trade-off claim is incomplete without it. 5. The reference distribution is fixed; reference staleness is unexamined. The protocol uses a fixed training-distribution reference for all 30 days. In production, the reference itself ages — pipelines change, populations shift — and deciding when to refresh the reference is one of the hardest operational questions in monitoring. A detector that is stable against a fixed reference may behave differently against a rolling one. Even a short discussion of how the findings interact with reference-refresh policies would help practitioners apply them. 6. The alerting-policy layer is missing. Practitioners do not page on raw detector output; they add deduplication windows, cooldowns, escalation thresholds, and multi-signal confirmation before an alert reaches a human. The paper's false-positive-day metric implicitly assumes one alarm = one unit of fatigue, but a well-designed alerting policy absorbs exactly the kind of fluctuation the statistical detectors show (0.2–0.6 FP days). Engaging with this layer — even briefly — would connect the findings to how monitoring is actually operated and sharpen the "cry wolf" claim. 7. Five seeds is thin for the headline threshold claim. The PSI ~200-sample transition is the paper's most actionable finding, and at batch 150 PSI shows 12.20±1.33 while batch 200 shows 1.60±1.02 — a sharp cliff where seed variance matters. More seeds around the transition region, plus a second dataset, would turn an observed threshold into a calibrated guideline. Relatedly, the paper should give readers a procedure for calibrating the minimum-batch-size threshold on their own data rather than a single number. Minor issues - MMD's near-zero TPR across all shift magnitudes (Table 8) deserves a deeper look — is this the default Gaussian bandwidth failing on a localized univariate shift? A bandwidth sensitivity check would tell readers whether MMD is truly unsuitable or just misconfigured. - The adversarial detector uses logistic regression with an AUC ≥ 0.6 alarm threshold; a stronger classifier might change its "conservative" characterization. The threshold choice (0.6) is asserted, not justified. - Time-to-detection is reported as N/A for zero-detection cases, which is fine, but the handling should be stated explicitly in §3.3 rather than left to the tables. - Figure 2's "upper-left quadrant is most desirable" framing is good; adding the Bonferroni variants as points on the same plot would visualize comment 4's trade-off directly. Overall assessment A valuable, practitioner-relevant paper on a real and under-studied problem — I have seen monitoring dashboards that engineers stopped looking at because the detectors cried wolf, and this paper measures exactly that failure mode. The operational metrics, the transparency of the experimental appendices, and the genuinely usable guidelines table are real strengths. My major comments ask for scoped generality claims, sounder categorical handling, compute-cost data, an FDR comparison, and engagement with the reference-refresh and alerting-policy layers that surround any real deployment. I would be glad to see this published once those are addressed. Competing interests The author declares that they have no competing interests. Use of Artificial Intelligence (AI) The author declares that they used generative AI to come up with new ideas for their review.

View source

Similar papers

#artificial intelligence Open access May 2023

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations...

Xiaotian Zhang, Chun-yan Li, Yi Zong et al. · 216 citations · ⚡17
#artificial intelligence Open access Jul 2024

Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval

This work investigates the possibilities of using LLMs in a resume screening setting via a document retrieval framework that simulates job candidate selection and finds that the MTEs are biased, significantly favoring White-associated names in 85% of cases and female-associated names in only 11.1% of cases.

Kyra Wilson, Aylin Caliskan · 131 citations · ⚡8
#artificial intelligence Review Oct 2025

Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition

This paper presents a comprehensive overview of the Ultralytics YOLO family, emphasizing architectural evolution, benchmarking, deployment, and emerging directions from YOLOv5 through YOLO27, and examines detection, segmentation, depth, classification, pose, oriented detection, tracking, export, quantization, and deplo...

Ranjan Sapkota, Manoj Karkee · 112 citations · ⚡10

BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models

A novel threat is unveiled in which attackers steer the RAG system's response by injecting malicious passages into its knowledge base, enabling the attacker to steer the response without altering the user input or modifying the RAG weights.

Jiaqi Xue, Meng Zheng, Yebowen Hu et al. · 109 citations · ⚡8

The Death of Schema Linking? Text-to-SQL in the Age of Well-Reasoned Language Models

This work revisits schema linking when using the latest generation of large language models (LLMs) and finds empirically that newer models are adept at utilizing relevant schema elements during generation even in the presence of large numbers of irrelevant ones.

Karime Maamari, Fadhil Abubaker, Daniel Jaroslawicz et al. · 109 citations · ⚡19

PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30% of conventional pipelines, and achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks.

Jinhe Bi, Yifan Wang, Danqi Yan et al. · 73 citations · ⚡4

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.