Human pairwise comparisons provide a reference for evaluating large language models (LLMs), but collecting sufficient judgments for each new release is costly and time-consuming. LLM judges offer a scalable alternative, although their comparisons may differ systematically from human preferences and across judges. We st...
Xin Zhou, Si-Nian Zhang, Zhan-Yan Yang et al.· 0 citations
GroupKV is presented, a lightweight hierarchical KV cache management system for long-context dLLM inference that observes that under block-wise decoding, tokens within the same generation block tend to access highly overlapping and spatially concentrated context regions, making group-level sparse selection effective.
Jin-Hao Wang, Zhe-Xin Hu, Kang-Jie Zhou et al.· Proceedings of the Internati...· 2 citations· ⚡1
Diffusion large language models (dLLMs) are emerging as a promising generative paradigm that complements autoregressive decoding. In long-context settings, KV cache bloat and offloading transfer overhead have become primary bottlenecks in inference systems. Meanwhile, the periodic full-sequence recomputation and locali...
Jin-Hao Wang, Zhe-Xin Hu, Kang-Jie Zhou et al.· Proceedings of the Internati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.