Modular Representation Compression: Adapting LLM Representations for Efficient and Effective Recommendation
Modular Representation Compression (MARC) is proposed to explicitly control the modularity of LLMs, and identifies a counterintuitive phenomenon during representation compression: Mid-layer Representation Advantage (MRA), where representations from middle layers of LLMs outperform those from final layers in recommendation tasks.