Editorial: Responsible and robust evaluation for real-world recommendation and search systems
Abstract
Both search and recommender systems are software tools that are used on a daily basis and influence a large number of decisions, affecting users, different stakeholders (companies, governments, organizations, etc.), and society as a whole (Ricci et al., 2022; Alonso and Baeza-Yates, 2024). Hence, as the evaluation of these systems is critical, it should go beyond performance prediction in traditional and/or artificial testing environments. Implemented real-world systems must deal with ambiguous data, temporal changes, conflicting interests among different stakeholders, and limits to generalization. All while they attempt to maintain stable performance despite changing circumstances.Therefore, the evaluation must be robust, i.e., examining whether performance and behavior remain consistent under variable real-world environments, and responsible, i.e., ensuring that the objectives and indicators reflect the users and environments impacted by the system. The four articles in this research topic explore some of these issues in four different domains: search query parsing, tourism recommender systems, career path prediction, and news moderation. Collectively, they show that the evaluation step cannot be viewed solely as the final stage of comparison between different proposals, but rather as a key component in the whole system design that reveals which system behaviors and trade-offs should be prioritized. Evaluation protocols should include realistic scenarios that examine potential failures that may occur under human supervision.In addition, assumptions and limitations of the studies must be clearly stated so that they can serve as the basis for responsible decision-making processes, showing the ambiguity and disagreements behind the overall scores.