Code and results for "Semantic Knowledge-Base Poisoning in AI-Native 6G: A Provenance-Based Meaning-Drift Audit for Trustworthy Federated Semantic Communication"
Abstract
This record contains the complete, self-contained code and precomputed results accompanying the article "Semantic Knowledge-Base Poisoning in AI-Native 6G: A Provenance-Based Meaning-Drift Audit for Trustworthy Federated Semantic Communication." It reproduces every table and figure in the paper. The study introduces semantic knowledge-base (KB) poisoning, a targeted attack on federated semantic communication in which a small set of malicious clients steer the recovered meaning of a chosen concept toward an attacker-selected concept while deliberately preserving classification accuracy, rendering the attack largely undetectable by accuracy monitoring. To detect and mitigate it, the framework anchors per-client semantic fingerprints of codebook contributions in a hash-chained ledger, producing a tamper-evident meaning-space drift audit; a simple threshold on this audit is shown to be the strongest, most robust defense. The software (package "s6g") is a reproducible testbed written in pure NumPy, with matplotlib for figures and scikit-learn for the real-modality dataset. It requires no GPU, no deep-learning framework, and no external downloads, and runs end to end on a single CPU in minutes; the transceiver's manual backpropagation is finite-difference verified. Included: the s6g package (synthetic and real semantic spaces, wireless channel, transceiver, attacks, aggregation and detection defenses, provenance ledger and drift audit, and a reinforcement-learning trust policy); experiment drivers for multi-seed statistics, real-modality validation on handwritten digits, an audit-aware adaptive adversary, and detection-based baselines (FLTrust, FoolsGold); raw per-seed results as JSON; generated figures; and a Word README documenting layout, commands, and results. Contents: README.docx, requirements.txt, the s6g/ package, four experiment driver scripts, results/ and figures/ directories, and *_results.json data files. To reproduce, install the requirements and run the drivers as documented in the README; each driver's "report" command prints the aggregated results. Keywords: 6G; semantic communication; federated learning; data poisoning; blockchain provenance; trust; Byzantine robustness.