Skip to content

Overview of the TREC 2025 Million Large Language Models track

Sep 2026 · 0 citations · 5 references
Computer Science

TL;DR

The TREC Million LLM Track operationalizes a retrieval-based paradigm in which an assistant agent infers expertise dynamically by examining models'observable behavior, providing the first large-scale benchmark for expertise retrieval in agentic AI.

Abstract

Agentic AI envisions ecosystems of intelligent agents collaboratively solving complex tasks with minimal human intervention. In such ecosystems, each agent possesses specialized expertise, making effective expert selection central to overall system performance. While most current approaches assume a small number of well-documented models, real-world expertise is far more diverse and cannot be adequately captured through static metadata or hand-written descriptions. We anticipate a future with millions of specialized language models (LLMs), each excelling in different domains or problem types. Rather than relying on predefined capability statements, we propose a retrieval-based paradigm in which an assistant agent infers expertise dynamically by examining models'observable behavior. Upon receiving a user query, the assistant ranks candidate LLMs based on demonstrated competence, enabling efficient and adaptive expert selection. The TREC Million LLM Track operationalizes this paradigm by shifting the retrieval target from documents to expert LLMs. Participants are given a discovery set consisting of queries, answers, and log-probabilities from more than one thousand LLMs and are challenged to infer meaningful expertise representations for each model. Given an unseen test query, systems must then rank the LLMs according to their expected performance, providing the first large-scale benchmark for expertise retrieval in agentic AI.

View source

Similar papers

Open access Sep 2026

Automating Plan Evaluation Using Agentic Large Language Models

This research benchmarks human evaluations against a large language model (LLM) using a multi-agent approach and/or retrieval-augmented generation (RAG) to automate complex content analysis tasks to leverage artificial intelligence’s efficiency and precision alongside humans’ contextual understanding and domain experti...

Xinyu Fu, Chaosu Li · 1 citation
Conference Open access Sep 2026

A Review on Test-Time Scaling for Agentic Large Language Models

A novel RAIE taxonomy along four scaling dimensions is proposed, which optimizes the entire thought process through search algorithms and self-verification, and introduces a task-oriented guideline for choosing the best TTS strategy.

Jia-Yu An, Zheng Chen, Yong-Cheng Jing et al. · 0 citations
#natural language process... Preprint Aug 2026

Improving LLM-based Autonomous Web Agents with Filtering

DeBERTa-based and T5-based models that rank HTML elements by their relevance to the task and a zero-shot ColBERT-based retriever that is able to retrieve the ground-truth element with a recall of 0.52 on Mind2Web and 0.47 on WebArena are developed.

Zhi-Tong Guo, Jing Yu Koh, Rui-Yu Li · 0 citations

Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers

Reproducibility is essential to scientific progress, yet the growing volume and complexity of scientific publications make exhaustive manual verification increasingly impractical. Although recent advances in large language model (LLM) agents enable automated experiment reproduction, existing evaluations largely focus o...

Han-Hua Hong, Yi-Zhi Li, H. Luu et al. · 1 citation
#artificial intelligence Preprint Oct 2026

AgentDiscover: Autonomous Discovery with Minimal Search Scaffolding

Frameworks that use large language models for scientific discovery typically rely on a fixed, human-designed algorithm that decides what the model sees at each step, leaving the model only the role of proposer. The model knows nothing of the search beyond what it is shown. As models grow more capable, a question arises...

Mahdi Farahbakhsh, I. Sela, Fatemeh Doudi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Schema: Discovering Unknown Environments via Agentic Program Induction

Learning to complete tasks in unfamiliar environments with unknown rules remains a key challenge for LLM agents. Current LLM agents often record their discoveries in prose, which may not provide a compact, explicit account of how the environment works. Inspired by how scientists organize observations into testable, pre...

Guan-Ning Zeng, Jia-Ni Wang, Wenjie Ma et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.