Shortcut-controlled performance attribution for multi-sensor acoustic emission source classification on a PMMA plate
Abstract
Multi-sensor acoustic emission (AE) source classification typically reconstructs multi-channel transient hits into graph-level samples to integrate transient waveforms with sensor topology information. However, the event reconstruction process may also introduce non-waveform structural confounders, such as trigger channel patterns and hit counts, making it difficult to determine the true source of high model performance. To address this issue, this study proposes a hierarchical shortcut-controlled evaluation framework for AE source classification on a PMMA plate, aimed at distinguishing waveform discriminability, graph-level modeling gain, and spurious correlations induced by reconstruction variables. The framework first evaluates the intrinsic separability of single waveforms at the hit level using continuous wavelet transform (CWT) features. At the graph level, it compares Graph-CWT-SVM, CNN-only, and a physically-informed weighted graph convolutional network (PI-WGCN), while incorporating two non-waveform baselines—event-count-only and channel-mask-only—as shortcut controls. In the experimental evaluation, hit-level CWT-SVM achieved a macro-F1 score of 0.964; in the standard graph-level assessment, Graph-CWT-SVM, CNN-only, and the masked PI-WGCN achieved 0.932, 0.923, and 0.917 macro-F1, respectively; the event-count-only and channel-mask-only baselines still reached 0.857 and 0.795 macro-F1. These results indicate that a substantial portion of the discriminative information in pseudo-event graphs can be explained by reconstruction variables rather than by waveform and topology-driven physical propagation features. This work highlights reconstruction-induced structural confounders in multi-sensor AE graph learning and provides a more rigorous, interpretable, and reproducible performance attribution approach for intelligent AE diagnostic models. While current data-driven methods have achieved high accuracy in AE source classification, the prevailing evaluation paradigm often overlooks this potential interference of structural biases. By addressing this critical oversight, our framework bridges the gap between idealized algorithmic metrics and actual waveform-driven diagnostic capabilities, establishing a more reliable evaluation standard for future research in the field.