Skip to content
Preprint

ARM: Detector-Agnostic Changepoint Attribution with Finite-Sample Error Control

Aug 2026 · 0 citations · 25 references
Mathematics Computer Science

TL;DR

ARM (Attribution by Rank Maxima), a wrapper that accepts a changepoint located by an arbitrary detector and returns the set of coordinates certified to have changed, each carrying a location or scale type label, is proposed.

Abstract

Detecting a change in a multivariate series answers only the first of two questions; the operational question is which coordinates changed. Existing answers are incomplete. Block-level procedures certify predefined groups of coordinates under an additive union bound, high-dimensional variable-selection methods return interpretable rankings without error guarantees, and the post-detection inference literature controls error along the time axis rather than across coordinates. We propose ARM (Attribution by Rank Maxima), a wrapper that accepts a changepoint located by an arbitrary detector and returns the set of coordinates certified to have changed, each carrying a location or scale type label. ARM scores each coordinate by a max-over-splits rank statistic. Because this statistic dominates the corresponding statistic at the estimated split, the resulting certificate is invariant to the manner, and to the accuracy, of the changepoint estimate. Three finite-sample guarantees follow from within-coordinate ranks alone: per-coordinate validity under any detector; exact family-wise error control through a Westfall--Young joint permutation that preserves cross-coordinate dependence, with a fully distribution-free Holm fallback; and false discovery rate control under arbitrary coordinate dependence in high dimensions through Benjamini--Yekutieli and e-BH. In simulations, naive per-coordinate testing at the estimated changepoint inflates its family-wise error beyond $0.66$ as the dimension grows, whereas ARM maintains the nominal level while retaining validity under heavy tails, power in high dimensions, and accurate type labels. On five financial series surrounding the 2008 collapse, ARM attributes a scale change to every asset class and excludes injected control coordinates.

View source

Similar papers

Preprint Aug 2026

Evidence, Calibration, and Stability: A Triadic Framework for Hypothesis Testing Under Model Uncertainty

Statistical tests are often asked to do too much. A single reported result is expected to describe what the observed data say, reassure readers about repeated-sampling behavior, and remain convincing when the working model is perturbed. Those tasks are connected, but they are not equivalent. Fisherian inductive inference and Neyman-Pearson decision theory clarify the first two; robust testing, sensitivity analysis, fragility measures, multiverse analysis, and distributional-stability methods speak to the third. I propose Evidence-Calibration-Stability (ECS) as a framework for keeping these roles separate while reporting them together. Evidence is post-data. Calibration belongs to the design or procedure. Stability is the post-data distance from the benchmark analysis to a conclusion-reversing perturbation within a declared model neighborhood. Full ECS support is conjunctive: a strong coordinate cannot rescue a failed one. For finite-dimensional affine perturbations, I derive an exact ellipsoidal stability radius. For smooth nonlinear margins, a uniform quadratic-remainder condition yields a certified lower bound over a declared neighborhood, showing when the affine formula is only a surrogate. I also establish coordinate invariance and a matrix extension for multiple claims, and distinguish confirmatory calibration from descriptive calibration profiles when prespecification is unavailable. Simulations for the one-sample t test and Student's historical sleep data show that the three coordinates can lead to different interpretations. ECS is a formal synthesis, not a claim that evidence, power, or robustness is itself new.

Subir Hait · 0 citations
Preprint Jul 2026

Identification and Inference with Machine-Learned Instruments

Instrumental-variables estimation increasingly pools many or high-dimensional instruments into a single machine-learned first stage, with rich controls partialled out. The resulting estimand, the partialled-out IV coefficient built from any signal of the instruments, is a signal-weighted average of the heterogeneous effects, which gives an opaque first stage a precise structural meaning. The average is convex whenever a covariance-monotonicity condition holds, and we provide a microfoundation for that condition based on vector monotonicity. With a learned signal, however, the usual debiased moment is not Neyman-orthogonal, and its first-order bias is a drift toward the learner's own signal-weighted average, so naive inference remains valid only for that learner-dependent target. We construct a heterogeneity-robust orthogonal score that restores $\sqrt{N}$ inference on the fixed, learner-invariant target at no efficiency cost, and provide a Hausman-type diagnostic and identification-robust confidence sets.

Fang Yu · 0 citations
Preprint Aug 2026

Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection

It is shown the natural way to do this does not work, specify one that survives measurement, then finds that the correction making it work carries more variance than the null it is tested against, and that the correction making it work carries more variance than the null it is tested against.

Floriane C. M. Braun · 0 citations
Preprint Aug 2026

Counterfactual Evaluation of Temporal Observation Protocols

We study counterfactual protocol evaluation: whether data collected under a realised observation protocol determine the predictive value of alternatives that were never deployed. Protocol value is the population $R^2$ of the Bayes-optimal predictor of a fixed trajectory-level target from the measurements an alternative would collect. We show that even infinite benchmark data need not determine this value: distinct latent covariance structures can induce the same benchmark measurement--target law while assigning different values to the same alternative. We develop a value-specific identification theory in which only latent ambiguity that changes the alternative's value matters. For linear targets, invisible covariance directions certify non-identification, while targeted measurements can restore identification without recovering the full latent covariance; an exact permutation construction extends the result to nonlinear aggregate targets. With finite dense calibration data, uniform error bounds control protocol-selection regret and distinguishable value gaps. Exact marginal gains then support cost-constrained, target-aware observation design. Simulations and retrospective analyses of Sleep-EDF and Long-Term AF show that broad temporal-layout differences can be more reliably distinguished than fine placements selected from finite data. Together, these results connect identification, calibration resolution and observation design for undeployed protocols.

Xi-Zhe Zhang · 0 citations
Preprint Aug 2026

Exact Inference in Fixed-Effect Regressions with Concentrated Identifying Variation

In fixed-effect regressions with many groups, fixed effects can absorb most identifying variation, leaving a handful of observations to carry what remains. When variation is this concentrated, conventional $t$-tests can reject a true null more than half the time, and any fixed critical value is either invalid or so conservative it has essentially no power. This paper builds an exact test from the design alone. A \textit{nuisance-annihilating contrast} is a linear combination of the treatment and fixed-effect dummies that eliminates the fixed effects without touching the outcome; sign-flipping these contrasts is then an exact symmetry of the null distribution at every sample size, under arbitrary heteroskedasticity. In two-way designs --- worker-firm, firm-time --- these contrasts are exactly the cycles of the bipartite mobility graph, so the movement that identifies the treatment effect is what makes exact inference possible. Exactness costs power: relative to an oracle test, a chosen set of cycles has an observable \textit{capture ratio} $\kap\in[0,1]$ and standard-error premium $\kap^{-1/2}$, and a packing algorithm resolves the capture-granularity trade-off. In the Grunfeld investment regression (single-observation score concentration $73.9\%$), 32 cycle contrasts capture $\kap=0.627$ of the identifying variation, giving an exact $95\%$ confidence interval of $[0.150,\,0.450]$.

S. Halkiewicz · 0 citations
Preprint Jul 2026

Robust Instrumental Variables: Sharp Rates and Inference under Adversarial Contamination

Because 2SLS is built from sample averages, a small number of observations can have a disproportionate effect on estimates and inference. We introduce W-2SLS, a simple drop-in robustification that replaces these averages by quantile-winsorized means. We analyze W-2SLS under adversarial contamination, which permits both the identities and the reported values of the contaminated observations to depend on the realized clean sample and therefore accommodates targeted or strategic manipulation. Under finite $m$-th moments, W-2SLS attains the minimax-sharp rate $\eta_{n}^{1-\frac1m}+n^{-1/2}$, where $\eta_n$ is the fraction of observations that may be altered. Matching lower bounds identify the exact contamination thresholds for uniform consistency, root-$n$ estimation, and centered Gaussian inference with the same first-order law as clean-sample 2SLS. When $\sqrt{n}\eta_{n}^{1-\frac1m}\to 0$ robustness is first-order free. We also construct feasible heteroskedasticity-robust inference and a winsorized Anderson--Rubin test valid under weak identification and adversarial contamination. Finally, even without contamination, ordinary 2SLS can have poor uniform finite-sample concentration, whereas W-2SLS admits confidence-calibrated sub-Gaussian deviation guarantees.

A. B. Kock, David Preinerstorfer · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.