Skip to content

Author

Deke Guo

We have 2 of 7 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#large language models Book Open access Sep 2026

MIGServe: Layout-Aware Multi-Instance GPU Management for Efficient LLM Serving

MIGServe treats the physical layout of MIG instances as a first-class scheduling dimension through three techniques: buddy-aware partition placement, which preserves large contiguous free blocks by allocating next to existing occupied buddies; proactive pair-matching migration, which consolidates fragmented half-full buddy pairs off the critical path of inference.

Jian-Wen Chen, Yun-Kai Liang, Bin Gao et al. · 0 citations
Book Open access Aug 2026

Fine-Grained Energy Accounting in Production LLM Serving

KV (Key-Value) volume is introduced, a physically grounded metric that captures the spatiotemporal footprint of a request’s KV cache occupancy, and it is shown that energy per KV volume (EPV) provides a stable and reproducible signature for modeling serving energy.

Xianyi Yuan, Hanlong Liao, Kunming Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.