ProvDA: VGAE-Based Synthetic Provenance Data Augmentation for APT Detection
Abstract
Constructing high-quality provenance graph datasets for APT detection remains challenging because attack-related behaviors are sparse, highly imbalanced, and hidden within large-scale benign system activities. This paper proposes ProvDA, a quality-controlled data augmentation framework for multi-structured provenance graphs. ProvDA first constructs heterogeneous temporal provenance graphs from raw system events and extracts minority attack-related subgraphs from causal, temporal, and entity-level neighborhoods. It then employs a variational graph autoencoder to learn the latent distribution of minority provenance patterns and generate synthetic attack-related edges or subgraphs through latent-space sampling and interpolation. To ensure the validity of generated samples, ProvDA further incorporates provenance-aware causal repair, logic repair, operation-level filtering, confidence filtering, deduplication, and sample-weight control before merging synthetic data into the original training set. Experiments on six DARPA TC provenance cases show that, under the same GAT-based detector and unchanged real test sets, ProvDA improves the average F1 score from 49.19% to 89.55% and Precision from 41.15% to 88.49%, while achieving 94.56% AUC and 92.54% AP on average. Compared with existing augmentation methods, ProvDA delivers stronger and more stable downstream detection utility across different provenance sources and engagement settings. These results demonstrate that quality-controlled minority subgraph generation can effectively enrich scarce attack patterns while preserving the structural, temporal, and operational consistency of provenance data.