Skip to content
Book Open access

Which LLM to Fine-Tune? Agent-Driven Model Selection at Scale

Sep 2026 · Proceedings of the 20th ACM Conference on Recommender Systems · pp. 1673-1681 · 0 citations · 10 references

TL;DR

It is shown that model selection is a recommendation problem, and AgentRec, a multi-stage retrieval-and-ranking framework that progressively narrows hundreds of candidate models using increasingly expensive but more faithful evaluation signals, is introduced.

Abstract

Open-source model hubs now host over two million public AI models, yet teams building customer-facing AI systems must still determine which model to fine-tune for production deployment—a decision that shapes the quality, latency, and cost experienced by hundreds of millions of users. At Amazon, we spent over years of iterating on this process across multiple production use cases, where model selection remained manual, slow, and heavily biased toward a small set of familiar model families despite the rapidly expanding open-source ecosystem. We show that model selection is a recommendation problem, and introduce AgentRec, a multi-stage retrieval-and-ranking framework that progressively narrows hundreds of candidate models using increasingly expensive but more faithful evaluation signals. Across public benchmarks and Amazon production systems, AgentRec reduces model selection from multi scientist-weeks to couple unattended GPU-hours while matching or exceeding the quality of exhaustive manual exploration. Our results suggest that, for industrial teams deploying fine-tuned LLMs at scale, model selection can evolve from an ad-hoc bottleneck into a repeatable and continuously automated system for discovering high-quality models under real-world deployment constraints.

Read PDF

Similar papers

#artificial intelligence Review Sep 2026

RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leav...

Zheng-Yu Chen, Lin-Feng Liu, Hong Li et al. · 0 citations
#natural language process... Preprint Sep 2026

CORAL: An LLM-Native Harness for Production Recommender Systems

Production recommender systems shape what billions of people see, and sustaining their performance requires continual optimization: as content, user behavior, and upstream models shift, the choices governing retrieval, ranking, and serving must be revisited. Traditionally, human engineers test such changes through onli...

Muhammad Azhar, Yu-Hang Zhou, Gilbert Jiang et al. · 1 citation
#machine learning Preprint Sep 2026

AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework

Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language model that iteratively edits a training script and retains modifications that improve a hel...

Aparajith Chandran, Juwon Kim, Saurav Jha et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Agentic ML Exploration (A-MLE) for Ads Ranking

Modern industrial ads ranking stacks are increasingly bottlenecked not by model capacity or training compute, but by the throughput of human ML iteration - the cycles of research, implementation, training, debugging, evaluation, and launch required to surface a single statistically significant improvement. A typical ra...

Erwin Gao, Vinodh Kumar Sunkara, Jing Guan et al. · 0 citations
#natural language process... Preprint Sep 2026

You're Hired: Strategic Model Selection for LLM Collaboration

While multi-agent and model collaboration algorithms gain traction to combine the strengths of diverse Large Language Models (LLMs), existing systems remain bottlenecked on pre-defined and hand-crafted model pools. In this work, we investigate the problem of model selection in multi-LLM systems. We propose and systemat...

Zong-Wan Cao, Zi-Yuan Yang, Shang-Bin Feng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.