An Empirical Analysis of Retrieval and Evaluation in the RecSys Challenge 2026 Music-CRS Track
Abstract
We describe a solo entry to the ACM RecSys Challenge 2026 Music-CRS track (team Siwon, CodaBench swlee9087): a frozen, inference-only system that combines BM25, a frozen instruction-tuned text embedder, and a listening-history centroid through Reciprocal Rank Fusion, followed by a two-stage response generator. It reached a composite of 0.3542 on the Blind B leaderboard (13th of 17 teams in the academic track, 33rd of 40 across both tracks) and 0.1853 on the interim Blind A board. The system is ordinary, so what we report is what we measured while building it, most of it negative. Retrieval quality drops steeply with dialogue depth in two pipelines that share no components; because the blind server scores only the deepest turn, an all-turn development average overstates blind-protocol retrieval by 1.7 to 2.0 times. The obvious fix, rebuilding the query around the pending request, does not help, and BM25 beats every dense query builder we tried; the dataset’s provided audio and collaborative-filtering embeddings add little zero-shot. The scoring itself carries measurable noise—two byte-identical prediction files drew different LLM-judge ratings—and score changes across the two blind phases confound data with system changes and cannot be read as generalization. Code, configurations and logs: https://github.com/swlee9087/acmrs2026_challenge.