Preprint
Aug 2026
When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide
This work benchmarks six estimators across five datasets and two known-effect sweeps, and validate the mechanisms against a non-simulated paired reference, finding that weak overlap is governed by logger-target action alignment, not by logging sharpness alone.
Binshuang Li
· 0 citations