Adaptive Detection of Unknown Threats in Smart Contracts via Multimodal Self-Learning
Abstract
Existing AI-assisted smart contract vulnerability detection methods train detection models with known vulnerability samples. These models predominantly emphasize unilateral feature representations of smart contracts, which impedes their ability to capture vulnerabilities comprehensively and undermines their generalization to previously unseen threats. This paper proposes Synesthete, a new method for detecting unknown smart contract threats. Synesthete adopts multiple types of smart contract features (e.g., texts, graph structures, and images extracted from opcodes, sources, and bytecode streams) and fuses the multimodal features to improve its cross-domain adaptability. Leveraging self-learning and data generation, Synesthete can cope with the imbalance of smart contract datasets (i.e., with relatively few vulnerability samples). By further introducing a domain-adaptive learning mechanism, Synesthete learns features shared across source and target domains (i.e., known/unknown threats). Evaluations show that compared with unimodal features, Synesthete’s self-learning fusion of multimodal features, combined with domain-adaptive learning on balanced datasets, improves performance significantly. Compared with the latest classical and AI-based methods, Synesthete’s detection accuracy on known reentrancy, timestamp dependence, and transaction-ordering dependence vulnerabilities is improved by 5.46%–28.78%. On discriminating unknown threats to smart contracts, Synesthete also achieves a high detection accuracy while outperforming existing deep-learning-based approaches.