GroupMask is proposed, which generates the group selectors of all layers with a lightweight hypernetwork, relaxes them with a Gumbel-Sigmoid parameterization and a straight-through estimator, and learns them through sparsity-budget regularization and self-distillation while keeping the pretrained weights frozen.
Zhen-Gao Li, Shuo-Qiu Li, Xiao-Fan Zhang et al.· 0 citations
This work proposes MoEGen, an adaptation framework that shifts MoE-based PEFT from expert selection to expert-conditioned parameter generation, and decouples expert capacity from adapter storage while enabling instance-conditioned adaptation.
Yiming Zeng, Lei Lu, Zexin Li et al.· 1 citation· ⚡1
SALT, a model-agnostic extractive framework that organizes per-sentence keywords into a trie ordered by sentence frequency (SF), a lightweight, reusable proxy for document thematic structure, reduces the prefill computation and memory cost of long-context prompts while remaining composable with KV-cache methods that ta...
MUGEN is proposed, a unified motion--language framework that pays neither cost: no codebook, one draw, and surpasses the discrete-token state of the art on every retrieval and alignment metric on SnapMoGen.
Zhankai Ye, Yukai Jin, Bingyang Wei et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.