A participant-level estimand is defined and an analysis-matched null is obtained by permuting participant labels and rerunning that entire workflow, and every step that reads labels belongs inside the permuted analysis.
Abstract
Longitudinal sensing studies routinely collect thousands of windows from a few dozen participants. The records are numerous; the independent scientific units are not. When the outcome is defined per participant, this mismatch makes apparently precise findings vulnerable to pseudo-replication, to partition choice, and to the ordinary analytic flexibility of comparing several pipelines before reporting one. Splitting on participants prevents a person's records from straddling a split, but it does not calibrate the label-dependent workflow fold construction, preprocessing, tuning, calibration, and candidate selection that produced the reported number. We define a participant-level estimand and obtain an analysis-matched null by permuting participant labels and rerunning that entire workflow. In controlled simulation, window-level inference rejects in 70-80% of replicates when no effect exists and a window bootstrap rejects at the same rate; a participant bootstrap still rejects at 10-17%; the analysis-matched test holds 0.025-0.100 across cohorts of 20 to 80 participants. Freezing the selected pipeline instead of repeating the search inflates Type-I error to 0.240 with eight candidates, where repeating it holds 0.040. Applied to two public cohorts, wrist actigraphy (n=55) yields participant AUROC 0.928 with p=0.0050, a conclusion that persists under a scale-robust rank-pooled statistic and under a matched permutation null computed after excluding hospitalized participants (p=0.0089). Smartphone sensing (n=38, 7 positives) yields 0.636 and does not reject (p=0.1724) despite sufficient resolution, with sensitivity 0.143. The practical rule is narrow: every step that reads labels belongs inside the permuted analysis, and repeated records do not create additional independent participants.
Electronic health records support a wide spectrum of clinical prediction and decision-support studies, but reproducible EHR research now requires more than training a single predictive model. As the field expands from machine learning and deep learning to LLM-based and agentic AI, differences in cohort construction, te...
Yinghao Zhu, Zi-Xiang Wang, Lei Gu et al.· Proceedings of the 32nd ACM...· 0 citations
A new inference method for conducting multiple-treatment comparisons involving endpoints within the generalized linear model (GLM) framework under covariate-adaptive randomization (CAR) that can effectively control Type I error while potentially improving power.
This work analyzes current benchmarking practices and introduces a novel decomposition framework that disentangles the contribution of distinct data-generating components, such as confounding, dose distribution non-uniformity, and response surface complexity, to estimator performance.
Christopher Bockel-Rickermann, Daan Caljon, Toon Vanderschueren et al.· Proceedings of the 32nd ACM...· 2 citations
When data are grouped, hierarchical or multilevel models are commonly used to account for group-level variation with group-specific parameters. Leave-one-group-out cross-validation (LOGO-CV) is a suitable tool for evaluating predictive performance for new groups, providing an estimator of the expected log predictive de...
A. Riha, Svenja Jedhoff, Paul-Christian Bürkner et al.· 0 citations
It is shown that a usable signal is available after convergence, when loss no longer distinguishes the two populations, and applying a fixed perturbation to a converged model's inputs flips the predictions of the latter far more often than the former.
This work proposes post-pretrained lasso selective inference (PPL-SI), a novel selective inference method designed to provide statistically valid p values for the pretrained lasso that reliably controls false discoveries and significantly improves the detection of biologically relevant features compared to traditional...
Cao Huyen My, Nguyen Vu Khai Tam, Vo Nguyen Le Duy· Statistics and computing· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.