An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic
This work forms model extraction monitoring as benign-calibrated traffic-window distribution testing: embed incoming queries into a semantic space and test whether their aggregate distribution deviates from historical benign traffic.