Skip to content

Beyond Heavy Log Curation: Perplexity-Based APT Detection via Unsupervised, Context-Augmented Language Models

Jul 2026 · arXiv.org · Vol abs/2607.20832 · 0 citations · 48 references
Computer Science

TL;DR

CAPTAIN (Context-Augmented Perplexity-based Threat Activity log detectIoN), a perplexity-based detector that leverages general, pre-trained language models with minimal, domain-agnostic preprocessing, enabling robust scoring of long, minimally processed log entries, is proposed.

Abstract

Advanced Persistent Threats (APTs) remain difficult to detect because only a small fraction of events in large-scale logs are attack-related, and investigation is expensive and hard to scale. Prior machine-learning approaches can reduce analyst workload, but they often rely on heavily curated training data and sophisticated preprocessing pipelines. Building and maintaining such pipelines require substantial domain expertise and engineering cost. Motivated by insights from a study of a strong APT detection baseline, we propose CAPTAIN (Context-Augmented Perplexity-based Threat Activity log detectIoN), a perplexity-based detector that leverages general, pre-trained language models with minimal, domain-agnostic preprocessing, enabling robust scoring of long, minimally processed log entries. CAPTAIN encodes recent history with an encoder model and a Q-Former-style bridge, then injects the compact context tokens into the decoder input so that perplexity reflects temporal context. To improve stability, CAPTAIN additionally applies smoothing filters to the perplexity time series. Across APT-oriented benchmarks, CAPTAIN competes with strong existing baselines and remains robust under substantially less curated inputs, that reduces the development and operational cost of advanced log preprocessing.

View source

Similar papers

#natural language process... Preprint Aug 2026

REER-PT: Reverse-Engineered Reasoning for Perplexity-Guided Pre-training Data Augmentation

Together, the perplexity analysis indicates improved continuation predictability, while the controlled pre-training experiments suggest that this augmentation can improve model performance without changing the standard pre-training objective.

Haoran Que, Jia-Jun Shi, Ting Huang et al. · 0 citations
Preprint Aug 2026

Linear Probing Provides Robust and Efficient Detection of Machine-Generated Text

This work analyses the linearity and quality of MGT representations and shows that simple linear probes outperform a wide range of detectors while being substantially more sample-efficient, and demonstrates the potential of linear probes as as robust and sample-efficient MGT detectors.

Gerrit Quaremba, Hanqi Yan, E. Black et al. · 0 citations

BERM: Low-Overhead Prompt-Injection Detection via In-Situ Benign Representation Modeling

BERM is introduced, a lightweight framework that performs in-situ detection by modeling a host LLM’s internal representations extracted during prefill, adding negligible overhead and reducing incremental inference overhead to near-zero.

Maihao Guo, Chaoyang Zhao, Jin-Qiao Wang · 0 citations
Jul 2026

RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention

In controlled evaluations at 32,768 tokens, RIS-Stochastic at 1% density and 70 ensemble seeds achieves 75.00% accuracy, outperforming the native dense baseline, demonstrating that sparse attention acts as a regularizer: low density over multiple seeds filters out sequence-level noise, whereas higher density reintroduces distractor noise.

A. R. Santos · 1 citation
Preprint Aug 2026

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

U N M ASK is presented, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation, and demonstrates that the discovery and validation stages generalize to reward model preference data.

Chidaksh Ravuru, Shashank Srivastava · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.