Information Design Against Gaming and Learning Adversaries
The Pareto frontier between the two defense objectives is characterized, and both rates are confirmed on seven binary-classification tasks spanning tabular, image, and language-model-feature inputs: label-plus-counterfactual access extracts the boundary with up to $200\times$ fewer queries than a published label-only b...