A VLM-based automated relevance evaluation pipeline deployed within Pinterest Search for online A/B experiments is presented, demonstrating that VLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency.
Abstract
Relevance evaluation plays a crucial role in personalized search systems, serving as a guardrail alongside user engagement metrics to ensure that search results align with user queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present a VLM-based automated relevance evaluation pipeline deployed within Pinterest Search for online A/B experiments. We rigorously validate the alignment between VLM-generated judgments and human annotations, demonstrating that VLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency. Leveraging VLM-based labeling further unlocks opportunities to expand the query set, optimize sampling design, and efficiently assess a wider range of search experiences at scale. This approach leads to higher-quality relevance metrics and significantly reduces the Minimum Detectable Effects (MDEs) in online experiment measurements.
AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.
Han Bao, Yue Huang, Yan-Bo Wang et al.· Proceedings of the 32nd ACM...· 0 citations
DiverValue-Bench is introduced, a population-aware benchmark for evaluating multi-dimensional value alignment across 74 countries/regions and it is shown that lightweight preference-based fine-tuning with Low-Rank Adaptation and Direct Preference Optimization substantially improves in-domain value alignment while yield...
Yao Liang, Dongcheng Zhao, Fei-Fei Zhao et al.· Proceedings of the Thirty-Fi...· 0 citations
Vision-language models (VLMs) increasingly power consumer-facing AI search, yet evaluating them on the diversity of everyday visual questions remains challenging. Existing benchmarks often target predefined capabilities, such as multi-hop retrieval or long-form synthesis, whereas users ask photo-grounded questions span...
Hao-Nan Jiang, Guo-Jian Zhan, Jian-Cong Xie et al.· 0 citations
Hybrid queries—natural language questions over structured data that require both database capabilities and LLM reasoning—have recently emerged as a prominent research topic. However, existing solutions remain overly dependent on manual workflows, and current benchmarks are limited in scale and diversity. To bridge this...
Bo Li, Chenzhan Wang, Long-Kang Lin et al.· Proceedings of the 32nd ACM...· 0 citations
A sequential behavioral alignment framework pairing fine-tuning with preference optimization over paired correct and counterfactual rationales is developed and applied, demonstrating that behavioral alignment mitigates bidirectional rationalization while delivering human-interpretable reasoning traces without manual pi...
Alireza S. Ziabari, Kat Ellis, Colleen E. Chan et al.· 0 citations
This work proposes prediction-powered evaluation, a framework that combines limited human judgments with large-scale automatic scores to obtain data-efficient system comparisons that are provably unbiased, and introduces the Prediction-Powered Saving Ratio (PPSR), a meta-metric that measures how much human annotation a...
Mingqi Gao, Anthony B. Sicilia, Weiye Shi· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.