Skip to content
Book Open access

POSTER: DeePCAP: Enabling High-Fidelity and Cost-Efficient Archival Packet Trace Storage

Aug 2026 · Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication · pp. 2268-2270 · 0 citations · 30 references
Computer Science

TL;DR

DeePCAP introduces a query-driven fidelity framework spanning packet- and flow-level queries to tackle the fidelity disconnection, and proposes a novel dimensionality reduction approach using frequency domain encoding to improve cost-fidelity trade-off.

Abstract

Long-term network packet traces (e.g., pcaps), if available, can enable and inform lots of management tasks. However, storing packet data at scale is very expensive, forcing operators to choose between coarse historical summaries or short retention windows. In this context, deep generative compression (DGC) offers a new hope to store compact model parameters and regenerate structurally accurate traces on demand. We evaluate the suitability of recent deep generative approaches [21, 24] for packet trace modeling and generation. We find that their fidelity metrics are disconnected from the domain-specific queries/use cases, and they have bad cost-fidelity trade-off. We propose DeePCAP, an end-to-end trace storage system to close this gap. DeePCAP introduces a query-driven fidelity framework spanning packet- and flow-level queries to tackle the fidelity disconnection, and proposes a novel dimensionality reduction approach using frequency domain encoding to improve cost-fidelity trade-off. Our preliminary results show that DeePCAP achieves the best fidelity on the 100+ query suite and the strongest cost-fidelity trade-off.

Read PDF

Similar papers

#small language model Preprint Aug 2026

ProxyFormer: A Dual-Stream Proxy Architecture for Ultra-Long Context and High-Resolution Generation

This work proposes ProxyFormer, a general dual-stream architecture built upon proxy tokens, and introduces factorized multi-level compression/decompression, layer-wise dynamic compression ratios, asymmetric dual embeddings, and a proxy-only KV-cache inference scheme.

Zhongpan Tang · 0 citations
Preprint Aug 2026

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

This paper introduces a mean correction term that effectively suppresses the approximation error, keeping performance degradation manageable even at extreme sparsity levels, and redesigns the sparse attention operator with PackGQA memory access, warp specialization, and pingpong pipelining.

Qihang Fan, Huaibo Huang, Zhiying Wu et al. · 0 citations
Open access Aug 2026

Loss-Resilient Semantic Communication over Packet-Loss Networks at Extreme-Low Bandwidth

In extreme-low bandwidth network scenarios, generative semantic codecs have emerged as promising solutions to reduce bandwidth cost for visual communications. However, these learned codecs are usually optimized solely for compression efficiency and thus not robust against transmission errors. Corruptions due to packet-loss among these highly compact generative latent representations often cause more critical degradation in fidelity and realism, intensified by the severe error propagation across the latent contexts and multi-step decoding process. In this paper, we propose ResiGLC, a novel loss-resilient generative latent coding framework designed for robust semantic communication over extreme-low bandwidth packet-loss networks. Motivated by the inherent goal-consistency between generation and compression, we sufficiently exploit the impressive in-context predictive capabilities of language models. Integrated with the masked learning strategy, our model supports arbitrary context modeling of latent codes, which could mitigate the error propagation and handle unpredictable packet loss patterns. At the receiver, a progressive resilient decoding pipeline is presented, which leverages both the contextual relationship of the latent codes and the multi-modal semantic prior in the generative latent space, separately. By jointly optimizing toward both compression efficiency and packet-loss resilience, our proposed progressive decoding mechanism offers graceful performance when dealing with dynamic packet losses. Through extensive experimental evaluations, we establish that under packet-loss network conditions, ResiGLC can effectively improve the loss-resilience in terms of perceptual fidelity and realism qualities with extreme-low bandwidth cost.

Shengshi Yao, Jincheng Dai, Sixian Wang et al. · 0 citations
Jul 2026

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

This paper introduces training-free Sol-Attn (Sparsifying online attention), which unifies dynamic routing, sparse computation, and approximation correction in a single online-softmax pass, achieving a better accuracy-efficiency trade-off in sparse attention.

Haopeng Li, Yitong Li, Junsong Chen et al. · 2 citations
Preprint Aug 2026

Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

The aim is a recipe deployable without a multi-week hyper-parameter search, which distills the 4-bit student directly from the original, uncompressed model, and is released open-weight as Hypernova-60B.

Bakbergen Ryskulov, Iker Garc'ia-Ferrero, David Montero et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding

The development of long-context Large Language Models (LLMs) is constrained by the memory bandwidth bottleneck and quadratic complexity of the attention mechanism during decoding. To overcome the inherent trade-offs between the memory overhead of metadata-based metrics and the computational inefficiency of adaptive selection strategies, we present Faster Flash Decoding (FFD), a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding. FFD integrates the selector and computer into a fully fused kernel, replacing external metadata indices with content-aware scanning via low-bit quantization. Furthermore, we introduce the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization. Offering a training-free and plug-and-play solution, FFD also enables the reuse of scanning results for computation, achieving up to 11.6x kernel-level speedup and scaling to 256K context length, with 2.37x end-to-end throughput improvement. Empirical validation on RULER and LongBench confirms that FFD maintains model accuracy while delivering high-ratio sparsity, with code available at https://github.com/qluoluo/faster-flash-decoding

Zhigeng Liu, Zhiyuan Ning, Ruixiao Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.