Skip to content
Preprint

Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search

Aug 2026 · 0 citations · 27 references
Computer Science

TL;DR

A VLM-based automated relevance evaluation pipeline deployed within Pinterest Search for online A/B experiments is presented, demonstrating that VLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency.

Abstract

Relevance evaluation plays a crucial role in personalized search systems, serving as a guardrail alongside user engagement metrics to ensure that search results align with user queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present a VLM-based automated relevance evaluation pipeline deployed within Pinterest Search for online A/B experiments. We rigorously validate the alignment between VLM-generated judgments and human annotations, demonstrating that VLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency. Leveraging VLM-based labeling further unlocks opportunities to expand the query set, optimize sampling design, and efficiently assess a wider range of search experiences at scale. This approach leads to higher-quality relevance metrics and significantly reduces the Minimum Detectable Effects (MDEs) in online experiment measurements.

View source

Similar papers

Book Open access Aug 2026

AutoDavis: Automatic and Dynamic Evaluation Protocol of Large Vision-Language Models on Visual Question-Answering

AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.

Han Bao, Yue Huang, Yan-Bo Wang et al. · 0 citations
Conference Open access Sep 2025

DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values

DiverValue-Bench is introduced, a population-aware benchmark for evaluating multi-dimensional value alignment across 74 countries/regions and it is shown that lightweight preference-based fine-tuning with Low-Rank Adaptation and Direct Preference Optimization substantially improves in-domain value alignment while yield...

Yao Liang, Dongcheng Zhao, Fei-Fei Zhao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities

Vision-language models (VLMs) increasingly power consumer-facing AI search, yet evaluating them on the diversity of everyday visual questions remains challenging. Existing benchmarks often target predefined capabilities, such as multi-hop retrieval or long-form synthesis, whereas users ask photo-grounded questions span...

Hao-Nan Jiang, Guo-Jian Zhan, Jian-Cong Xie et al. · 0 citations
Book Open access Aug 2026

Automating End-to-End Hybrid Query Processing: Benchmark, Solution, and Insights

Hybrid queries—natural language questions over structured data that require both database capabilities and LLM reasoning—have recently emerged as a prominent research topic. However, existing solutions remain overly dependent on manual workflows, and current benchmarks are limited in scale and diversity. To bridge this...

Bo Li, Chenzhan Wang, Long-Kang Lin et al. · 0 citations
Preprint Aug 2026

From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation

A sequential behavioral alignment framework pairing fine-tuning with preference optimization over paired correct and counterfactual rationales is developed and applied, demonstrating that behavioral alignment mitigates bidirectional rationalization while delivering human-interpretable reasoning traces without manual pi...

Alireza S. Ziabari, Kat Ellis, Colleen E. Chan et al. · 0 citations
#natural language process... Preprint Aug 2026

Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

This work proposes prediction-powered evaluation, a framework that combines limited human judgments with large-scale automatic scores to obtain data-efficient system comparisons that are provably unbiased, and introduces the Prediction-Powered Saving Ratio (PPSR), a meta-metric that measures how much human annotation a...

Mingqi Gao, Anthony B. Sicilia, Weiye Shi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.