Choosing What Matters: Query-Aware Multimodal Routing for Conversational Recommendation
Abstract
Conversational recommender systems can exploit structured, collaborative, textual, and visual evidence, but fixed multimodal fusion cannot adapt evidence use to each dialogue. Building on MSCRS [24], we propose Query-Aware Multimodal Routing for Conversational Recommendation (QAMR-CRS), which predicts modality weights from dialogue context and mentioned entities, routes knowledge graph, co-occurrence, text-similarity, and image-similarity evidence at the entity level, and converts the routed evidence into prompt prefixes. In controlled 10-seed experiments, QAMR-CRS significantly improves Recall@10, Recall@50, NDCG@10, and NDCG@50 on ReDial and Recall@50 on INSPIRED over matched MSCRS reproductions. Ablations attribute the principal gains to query-dependent routing, while routing analysis shows that validation-selected residual calibration mitigates KG-dominant modality concentration with a modest ranking trade-off. Code is available at https://github.com/HyelimPark77/QAMR-CRS.