Code language models must be maintained like the software around them: when a library evolves, a model keeps writing the interface that it saw during training. Repairing the model itself lets one correction reach all downstream uses. Existing repair methods attribute a failure to neurons, select the highest-ranked ones...
Jian Gu, Hong-Yu Zhang, Chun-Yang Chen et al.· 0 citations
REFLEX is proposed, a training-free method that keeps the default router unchanged while reorganizing expert computation around the evolving refinement process, and introduces a coarse-to-fine hierarchy for expert-budget allocation that aligns computation with block-relative refinement roles while using the Frontier-Pr...
Xiang-Wen Xia, Chen Yan, Yiming Zhang et al.· 0 citations
RoleMerge, a training-free method that constructs each expert's Routing Role Profile (RRP) from phase-normalized routing statistics, capturing its relative phase preference, is proposed, and results validate phase-conditioned expert roles as a more effective basis than global routing aggregation for MoE-VLM expert merg...
Hong-Yu Zhang, Cheng Yan, Xiang-Wen Xia et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.