Skip to content

Author

Bin Gao

We have 2 of 9 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#large language models Book Open access Sep 2026

MIGServe: Layout-Aware Multi-Instance GPU Management for Efficient LLM Serving

MIGServe treats the physical layout of MIG instances as a first-class scheduling dimension through three techniques: buddy-aware partition placement, which preserves large contiguous free blocks by allocating next to existing occupied buddies; proactive pair-matching migration, which consolidates fragmented half-full buddy pairs off the critical path of inference.

Jian-Wen Chen, Yun-Kai Liang, Bin Gao et al. · 0 citations
Preprint Aug 2026

SAEM: Stage-Aware Expert Management for Memory-Efficient MoE Inference in Chain-of-Thought Reasoning

SAEM is proposed, a stage-aware MoE inference runtime that detects reasoning stage boundaries and exploits stage-level activation coherence to guide expert placement and achieves an average 1.33x throughput improvement over the strongest state-of-the-art caching and offloading baselines under constrained GPU memory.

Yujie Zhang, Bin Gao, Tulika Mitra · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.