OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit
OMP-MoE, a novel training-free compression framework for reducing expert redundancy in MoE-based LLMs by reformulate the pruning problem as a sparse signal reconstruction task solved through Orthogonal Matching Pursuit.