Skip to content

Not All Relations Are Equal: Relation-Balanced and Calibrated Graph Learning for Provenance-Based Intrusion Detection

Sep 2026 · 0 citations · 23 references
Computer Science

TL;DR

RECAL is presented, an unsupervised framework using relation-balanced masked graph learning to better capture rare interaction patterns and calibrates reconstruction errors against each relation's benign error distribution to produce comparable anomaly evidence, helping distinguish attacks from benign behavior and reduce false alarms.

Abstract

Provenance-Based Intrusion Detection Systems (PIDSs) detect Advanced Persistent Threats (APTs) by analyzing system interactions. However, existing methods largely treat relations uniformly, overlooking statistical heterogeneity; in CADETS, relation frequencies differ by approximately $140{,}000\times$. This may cause PIDSs to focus more on frequent relations and overlook differences in normal error levels across relations, increasing the risk of false alarms and missed detections. We present RECAL, an unsupervised framework using relation-balanced masked graph learning to better capture rare interaction patterns. It further calibrates reconstruction errors against each relation's benign error distribution to produce comparable anomaly evidence, helping distinguish attacks from benign behavior and reduce false alarms. On three DARPA E3 datasets, RECAL achieves F1 scores of 99.99\%, 99.93\%, and 99.99\%, outperforming the best baseline on each dataset by 0.88, 0.82, and 0.42 percentage points, respectively. Compared with the baseline reporting the lowest FPR, RECAL reduces mean FPR by approximately $105\times$, $4\times$, and $41\times$.

View source

Similar papers

Preprint Sep 2026

Where the Numbers Come From: Auditing Evaluation in Provenance-Based Intrusion Detection

Reproducing a provenance-based intrusion detector's score does not establish what that score says about its emitted alarms or the information its encoder uses. We audit nine released implementations, execute four detectors using their own code, and isolate three measurement effects. First, a fixed-alert comparison sepa...

Jih-Wan Moon, Gunhee Kim, Myeongjang Pyeon · 0 citations
Preprint Aug 2026

Provenance, Not Behaviour: A Serialisation Artifact in Edge-IIoTset and a Leakage-Free Benchmark for Precision-Agriculture Intrusion Detection

Edge-IIoTset is the reference benchmark for machine-learning intrusion detection in the industrial Internet of Things, and results reported on it cluster above 99%. We show that much of that performance is not intrusion detection. The preprocessing recipe distributed with the dataset instructs researchers to one-hot en...

M. Galal · 0 citations
Open access Sep 2026

Flow Exporter Provenance as a Major Confounder in Cross-Dataset IoT Intrusion Detection

Machine-learning intrusion detection for the Internet of Things (IoT) routinely exceeds 99% accuracy on single datasets but fails when transferred to new networks or flow exporters. We formalize five failure modes, define a 23-feature canonical schema, and adapt three datasets (CICIoT2023, TON_IoT, Bot-IoT) to build a...

Murad A. Rassam, Mahfoudh Alasaly · 0 citations
2026

CrossHound: Multi-Source Provenance Correlation for Advanced Persistent Threat Detection and Investigation

Advanced Persistent Threat (APT) attacks pose severe cybersecurity challenges due to their stealthy, prolonged, and highly targeted nature. Despite advancements in defensive technologies, detecting APTs remains difficult because their attack traces are often scattered across multiple data sources, leading to fragmented...

Jia-Ping Gui, Yi-Xin Zhou, Xiang-Geng Zhu et al. · 0 citations
Preprint Aug 2026

Amortised Post-Hoc Explanation with Exact Preservation for Dynamic Graph Anomaly Detectors

Anomaly detection in dynamic graphs underpins financial fraud analysis, intrusion detection, and platform integrity, where automated decisions require human-interpretable justifications. StrGNN, the strongest performer in recent benchmarks, produces no explanation: when an edge is flagged, the analyst receives only a s...

Iyad Assaad Nekka, H. Seba, Walid Khaled Hidouci et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.