{"slug": "onebid-a-unified-auto-bidding-foundation-model-for-diverse-ocpx-advertising", "title": "OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios", "summary": "Kuaishou deployed OneBid, a unified auto-bidding foundation model built on Decision Transformer, reporting an overall +2.2% ADVV gain on oCPX Ads and a peak +13.1% gain in the ROAS scenario in online A/B tests. OneBid extends single Return-to-Go conditioning to Return-to-Go for conversion value and Cost-to-Go for cost ratio, adds value-aware regularization on next-action prediction, and uses a sequence-level Mixture-of-Experts architecture with shared and sparsely-routed experts to handle heterogeneous oCPX scenarios at low latency. Post-training aligns the backbone to scenario preferences via Critic-guided Relative Offline Policy optimization (CROP), which avoids the unsafe online exploration of GRPO-style fine-tuning.", "body_md": "arXiv:2609.21550v1 Announce Type: new \nAbstract: Auto-bidding is central to computational advertising, where strategies must maximize advertisers' conversion value under economic constraints. It has evolved from rule-based controllers to reinforcement learning and generative methods such as Decision Transformer (DT). Yet these methods increasingly mismatch the prevailing optimized cost-per-X (oCPX) paradigm, which spans heterogeneous scenarios (e.g., registration, purchase), each served by a separate model, leading to fragmented pipelines and underexploring cross-scenario modeling. Inspired by foundation models like LLMs, unifying these oCPX scenarios into one model raises three challenges: multi-objective control, scalable capacity under strict latency, and safe offline policy improvement. We present OneBid, a unified auto-bidding foundation model that learns a reusable backbone from heterogeneous oCPX logs and adapts it to scenario-specific deployments via offline post-training. Building on DT, OneBid extends single Return-to-Go conditioning to two atomic signals, Return-to-Go for conversion value and Cost-to-Go for cost ratio, plus value-aware regularization on next-action prediction. To absorb distributional heterogeneity, we design a sequence-level Mixture-of-Experts architecture, where shared experts encode cross-scenario knowledge and sparsely-routed experts capture scenario-specific patterns at low latency, yielding consistent scaling with model size and data. During post-training, we align the backbone with scenario preferences via Critic-guided Relative Offline Policy optimization (CROP): a learned critic scores candidate actions group-relatively, avoiding the unsafe online exploration of GRPO-style fine-tuning while constraining policy shift to reduce OOD risk. Validated via online A/B tests and fully deployed at Kuaishou, OneBid delivers an overall +2.2% ADVV gain on oCPX Ads, peaking at +13.1% in the ROAS scenario.", "url": "https://wpnews.pro/news/onebid-a-unified-auto-bidding-foundation-model-for-diverse-ocpx-advertising", "canonical_source": "https://www.machinebrief.com/news/onebid-a-unified-auto-bidding-foundation-model-for-diverse-o-wwvf", "published_at": "2026-09-21 04:00:00+00:00", "updated_at": "2026-09-21 05:54:12.764823+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "artificial-intelligence"], "entities": ["OneBid", "Kuaishou", "Decision Transformer", "Mixture-of-Experts", "Critic-guided Relative Offline Policy optimization", "GRPO", "oCPX"], "alternates": {"html": "https://wpnews.pro/news/onebid-a-unified-auto-bidding-foundation-model-for-diverse-ocpx-advertising", "markdown": "https://wpnews.pro/news/onebid-a-unified-auto-bidding-foundation-model-for-diverse-ocpx-advertising.md", "text": "https://wpnews.pro/news/onebid-a-unified-auto-bidding-foundation-model-for-diverse-ocpx-advertising.txt", "jsonld": "https://wpnews.pro/news/onebid-a-unified-auto-bidding-foundation-model-for-diverse-ocpx-advertising.jsonld"}}