Skip to content

Author

Quang-Dung Dang

We have 2 of 5 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

DynaCalKV: Key-Value Cache Compression via Head Grouping and Adaptive Rank Allocation

As the inference phase of Large Language Models (LLMs) requires handling long context windows, the Key-Value (KV) cache initially appears to address this challenge but eventually becomes a significant bottleneck as the context window continues to grow. Low-rank compression has recently been studied as an effective approach to reduce KV cache memory while maintaining model performance. However, only a few existing methods treat the Key and Value caches differently, despite their distinct roles. Moreover, these methods typically employ fixed attention-head grouping, which may not fully exploit the structural similarity among attention heads. In this paper, we propose an improved low-rank KV cache compression framework. For the Key cache, we dynamically group attention heads based on Centered Kernel Alignment (CKA) similarity and allocate the rank budget adaptively under a parameter budget. For the Value cache, we adopt the same approach as ReCalKV, refining the low-rank decomposition through offline calibration to improve reconstruction quality. Experimental results on three instruction-tuned LLMs show that our method reduces the number of Key cache parameters while maintaining competitive accuracy. We further observe that the proposed strategy is particularly effective for Multi-Head Attention (MHA) models, whereas it should be applied more conservatively to Grouped-Query Attention (GQA) models, especially in long-context settings.

T. T. Nguyen, Quang-Dung Dang · 0 citations
Open access 2026

Persistent Goal-Tracking and Instruction-Driven Reasoning in Sequential Conversation QA

GSC-QA (Goal-based Sequential Conversation QA), a framework that integrates three complementary components into a unified enterprise dialogue architecture that combines retrieval, instruction enforcement, and goal persistence in a single coordinated loop built on LangGraph, is introduced.

Quoc-Dung Ngo, Quang-Dung Dang, Ly-Huynh Phan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.