Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 5674-5685· 0 citations· 12 references
Abstract
Labeling time-series events such as anomalies and system failures is expensive and subjective, and distribution drift makes label definitions evolve over time. Large language models can generate weak labels with natural-language rationales, but their error rates are unbounded and their failure modes opaque. We present CALM-TS, a risk-controlled weak-supervision pipeline that bounds the labeling error rate while maximizing coverage. CALM-TS introduces behavioral probing over prompt surface form, sampling temperature, and temporal context window to elicit disagreement signals; a lightweight calibrator converts them into risk-bounded acceptance with finite-sample guarantees. CALM-TS attains 77--81% cost reduction on MIMIC-III and Yahoo~S5 at empirical risk ≤ a=0.05 in all 10 dataset-seed pairs. On the official PhysioNet Challenge 2015 binary alarm-verification task with gpt-4o-mini over five seeds, CALM-TS is the only method among nine LLM-as-weak-labeler baselines whose 5-seed mean risk strictly falls below the unfiltered LLM at non-trivial coverage, delivering a 15.0% relative reduction (0.400 → 0.340) at coverage 0.286. The framework yields auditable evidence chains and a coverage-based, label-free drift indicator.
Oil-well anomaly monitoring supports safe and efficient oil-and-gas production, but delayed recognition of abnormal operating states can reduce lifting efficiency, trigger costly interventions, and increase operational risk. Existing data-driven detectors are also vulnerable to optimistic estimates when segmentation, n...
Feng Ge, Zhi Yang, Fan Yu et al.· Processes· 0 citations
Early time-series classification (ETSC) aims to make accurate predictions from partially observed time series as early as possible. Although various stopping mechanisms and feature learning strategies have been developed for ETSC, most existing methods assume access to sufficient labeled training data, which may be unr...
Chen-An Tai, Yujia Wu, Vincent S. Tseng· 0 citations
It is shown that TTA can increase confidence and reduce entropy even when the top-1 prediction and its correctness remain unchanged, a failure mode the authors term prediction-preserving sharpening, and proposed Zero-Shot-Anchored Entropy Calibration (ZAEC), a label-free post-hoc method that uses zero-shot entropy as a...
Jing-Yan Jiang, Yaru Sun, Xiao Chen et al.· 0 citations
This work proposes distilling the model into a probabilistic classifier, enabling lightweight deployment without repeated LLM calls, and demonstrates that LSR improves macro-F1 scores by an average of 7.0% compared to standard zero-shot classification baselines.
Nathan Vandemoortele, Bram Steenwinckel, F. Ongenae et al.· Discover Computing· 0 citations
Test-time adaptation (TTA) offers many ways to update a deployed model without labels, but choosing the wrong update can make a strong source model worse. Recent methods therefore try to predict which adaptation will work from unlabeled test data. We ask a prior question: does the evidence given to the selector contain...
This work introduces Gradient Activation Adaptive Multi-Level Splitting (GA-AMLS), which adapts rare-event Monte Carlo methods to the continuous activation space of language models and establishes activation space as a tractable domain for rare-event estimation in language models, circumventing the brittleness of discr...
Nikita Y. Parulekar, Anqi Liu· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.