When the Label Ignores the Request: Auditing Policy-Selected Targets in Synthetic Conversational Music Recommendation
Synthetic dialogues generated by LLM pipelines now serve as complete conversational-recommendation benchmarks: an LLM listener talks to an LLM recommender, and the track logged next in the conversation becomes the official label for each turn. These policy-selected labels make large-scale evaluation reproducible, but t...