Skip to content

Author

Peng Jiang

We have 6 of 23 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Aug 2026

Generative Auto-Bidding in Large-Scale Auctions via Diffusion Completer-Aligner

This work proposes a Causal auto-Bidding method based on a Diffusion completer-aligner framework, termed CBD, which achieves superior performance on large-scale auto-bidding benchmarks, but also delivers significant improvements on an online advertising platform, including a 2.0% increase in target cost.

Ye-Wen Li, Jingtong Gao, Peng Jiang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios

Auto-bidding is central to computational advertising, where strategies must maximize advertisers'conversion value under economic constraints. It has evolved from rule-based controllers to reinforcement learning and generative methods such as Decision Transformer (DT). Yet these methods increasingly mismatch the prevail...

Ye-Wen Li, Peng Jiang, Yi-Tian Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories

AgentBrew is proposed, an offline training framework that learns effective tool-use policies from a single batch of raw interaction trajectories, without task verifiers or iterative on-policy rollouts, and demonstrates that fine-grained offline learning can recover useful supervision from raw trajectories that filterin...

Zhiyi Lyu, Ye-Wen Li, Long-Tao Zheng et al. · 2 citations
Jul 2026

From Understanding to Action: Feedback-Grounded Policy Discovery for Generative Recommendation

Semantic-ID-based generative recommenders enable efficient next-item generation, but their item-level supervision mainly captures behavioral co-occurrence and local transitions. Large language models (LLMs) can complement these models by reasoning over heterogeneous interaction histories to understand the user's curren...

Zhi Chen, Minmao Wang, Xing-Chen Liu et al. · 0 citations
Book Open access Aug 2026

Hierarchical Residual Policy Optimization for Generative Recommendations

Hierarchical Residual Policy Optimization (HRPO), a post-training framework that converts item-level outcomes into dense, token-aligned learning signals for conservative token-wise improvement, is proposed.

Kaifeng Guo, Yiming Yang, Jingtong Gao et al. · 0 citations
Conference Open access Sep 2026

Adversarial Attack Framework Against Vision-Language Model Unlearning

Large Vision--Language Models (VLMs) unlearning tends to eliminate the influence of ``to-be-forgotten'' content in the training corpora, algorithmically by suppressing the likelihood of faithfully generating responses on forget-target inputs. The injection of adversarial inputs can manipulate the unlearned VLM's genera...

Yi-Min Liu, Peng Jiang, Ya-Jie Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.