Skip to content
Book Open access

Generative Auto-Bidding in Large-Scale Auctions via Diffusion Completer-Aligner

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 7612-7623 · 0 citations · 63 references

TL;DR

This work proposes a Causal auto-Bidding method based on a Diffusion completer-aligner framework, termed CBD, which achieves superior performance on large-scale auto-bidding benchmarks, but also delivers significant improvements on an online advertising platform, including a 2.0% increase in target cost.

Abstract

Bid optimization strategy in auto-bidding is central to computational advertising, achieving notable commercial success by optimizing advertisers' bids within constraints. Recently, generative models have revolutionized auto-bidding by directly learning a policy from large-scale datasets. Among them, the diffuser is superior in tackling sparse-reward challenges, along with its trajectory stitching and explainability capabilities, making it well-suited for industrial auto-bidding. However, its performance could be limited by generation uncertainty, particularly regarding generations' dynamic illegitimacy and preference misalignment, which can lead to suboptimal bids and further cause poor performance when competing with other advertisers in highly competitive auctions. To address it, we propose a Causal auto-Bidding method based on a Diffusion completer-aligner framework, termed CBD. Firstly, we conduct a theoretical analysis and propose a completer to augment the training process with an extra random variable t representing the decision timestep for enhancing the dynamic legitimacy between adjacent states. Then, we employ a trajectory-level return model as an aligner to refine the generated trajectories in inference for better alignment with advertisers' objectives. Experiments across diverse settings demonstrate that our approach not only achieves superior performance on large-scale auto-bidding benchmarks, such as a 29.9% improvement of conversion value in the challenging sparse-reward setting, but also delivers significant improvements on an online advertising platform, including a 2.0% increase in target cost.

Read PDF

Similar papers

Preprint Sep 2026

Efficiency of Generalized Proportional First-Price Auctions Under Auto-bidding

Auto-bidding is now widely adopted in online advertising platforms, allowing advertisers to specify high-level campaign objectives--such as maximizing total value subject to a return-on-spend (ROS) constraint--rather than manual per-query bids. A central question in algorithmic mechanism design is characterizing the wo...

Yang Cai, Vineet Gupta, Yan-Chen Jiang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios

Auto-bidding is central to computational advertising, where strategies must maximize advertisers'conversion value under economic constraints. It has evolved from rule-based controllers to reinforcement learning and generative methods such as Decision Transformer (DT). Yet these methods increasingly mismatch the prevail...

Ye-Wen Li, Peng Jiang, Yi-Tian Li et al. · 0 citations
Preprint Aug 2026

Fine-Tuning Autobidders with Group Relative Policy Optimization

Automated bidding (autobidding) is a core component of modern online advertising systems. Within this component, advertisers delegate sequential bid decisions to algorithms that must maximize campaign value while adhering to constraints such as a limited budget and a target cost-per-click (CPC). One of the approaches t...

A. Safin, Alexandra Khirianova, Andrey Pudovikov et al. · 0 citations
Preprint Sep 2026

Two-sided Market Design Meets Autobidding

Autobidding has become a dominant paradigm in online advertising by enabling advertisers to set high-level goals while algorithms handle real-time bid optimization. A prominent example is the Return-on-Spend (RoS) value maximizer, which maximizes total value subject to an aggregate value-per-spend constraint. While mos...

Yang Cai, Christopher Liaw, Aranyak Mehta et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Learning to Allocate Incentives for Incentivized Advertising via Offline Model-Based Reinforcement Learning

An offline model-based RL framework for cost-controllable sequential incentive allocation is developed and an independent counterfactual scorer evaluates each learned policy on held-out logs, enabling pre-launch selection without costly online exposure.

Zi-Lin Zhao, Han Yang, Tian-Pei Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.