DO CAN INTRUSION DETECTORS TRAVEL? A CROSS-DATASET AND ZERO-DAY STUDY WITH IDENTIFIER-AGNOSTIC, LOCALLY CALIBRATED ENSEMBLES
Abstract
Intrusion detection systems (IDS) for the vehicular CAN bus that use Machine Learning are usually evaluated and reported to achieve very high detection accuracy, commonly above 99%, on random splits of a single dataset. We test and reproduce such high detection accuracy for four publicly available CAN IDS benchmark corpora (OTIDS, Car-Hacking, ROAD, SynCAN), including three real-world vehicles and a synthetic generator, and we expose this number as an artifact of the protocol evaluation used so far. We present a detailed evaluation using a highly effective random-forest detector in different settings, i.e., using different corpora for training and testing. Our results show a significant collapse of a detector’s mean F1 score from near-perfect in-domain performance to 0.345 when transferred to other corpora. Evidence-based identification of this collapse is a significant contribution of this work, a problem which we subsequently fixed by removing identifier one-hot features and using local statistics that are both identifier-agnostic and are locally calibrated on a short window of unlabelled benign traffic of the target-vehicle. We subsequently achieve a significant increase in mean cross-dataset F1 score without using any labelled attack on the target corpus. To the best of our knowledge, this is the first work that (i) identifies the root-cause of this collapse, namely the feature representation, (ii) presents and repairs this feature-representation to recover much of the lost cross-dataset performance using identifier-agnostic statistics that are locally calibrated on minutes of unlabelled benign target-vehicle traffic, and (iii) reports per-family zero-day recall for four publicly available corpora.