{"slug": "uplifting-conversion-across-the-acquisition-funnel-with-personalization-using-on", "title": "Uplifting conversion across the acquisition funnel with personalization using contextual bandits on AWS", "summary": "Amazon Payments applied a multi-objective contextual multi-armed bandit on Amazon SageMaker AI to personalize a product acquisition funnel, reporting a high single-digit percentage relative lift in final-funnel conversion for one customer population over a seven-week online A/B test while a second population saw no improvement. The team used the Upper Confidence Bound (UCB) selection strategy for auditable, reproducible per-impression decisions and published a code repository for testing the approach on synthetic data. The result indicated the limiting factor was the content rather than the model.", "body_md": "## [Artificial Intelligence](https://aws.amazon.com/blogs/machine-learning/)\n\n# Uplifting conversion across the acquisition funnel with personalization using contextual bandits on AWS\n\nGenerative AI has made it possible to produce large amounts of personalized content quickly and at low cost. In [our previous post](https://aws.amazon.com/blogs/machine-learning/reinvent-personalization-with-generative-ai-on-amazon-bedrock-using-task-decomposition-for-agentic-workflows/), we showed how generative AI on [Amazon Bedrock](https://aws.amazon.com/bedrock/) can produce personalized content at scale while staying within brand guidelines and guardrails. The new challenge is now one of selection. Among all of those options, which one do you show each customer, and how long does it take to learn the answer? This post tackles the selection challenge that follows the proliferation.\n\nIn this post, we share how Amazon Payments applied AI-based personalization to a product acquisition funnel, using a multi-objective contextual multi-armed bandit (MAB) on [Amazon SageMaker AI](https://aws.amazon.com/sagemaker/ai/). In a seven-week online A/B test we currently see a high single-digit percentage relative lift in final-funnel conversion for one customer population, while another saw no improvement over the existing experience. The problem turned out to be the content, not the model. We cover the intuition behind bandits, our extension to optimize an entire conversion funnel, and the AWS architecture behind the solution. We also share [a code repository](https://github.com/aws-samples/sample-multi-objective-contextual-bandit-with-sagemaker) that you can use to test this approach on synthetic data and understand the method hands-on using Amazon SageMaker AI.\n\n## Growing role of multi-armed bandits in the generative AI era\n\n[A multi-armed bandit (MAB)](https://en.wikipedia.org/wiki/Multi-armed_bandit) is a reinforcement learning method built for settings with many options and limited traffic. It learns which option performs best while continuing to serve customers. It treats each content variation as an “arm,” tries each against live traffic, and steadily shifts impressions toward the arms that perform, while holding a fraction back to keep testing the rest. This is the fundamental trade-off between exploitation (serve the current best arm) and exploration (try less-certain arms to gather evidence). Because a bandit never stops doing both, it keeps improving as new variations are added, and it never has to wait for a test to conclude. A/B/n testing still has its place, but as generative AI accelerates the number of variations to learn from, we expect bandits to play a growing role.\n\nThere are several bandit selection strategies in the literature: epsilon-greedy, Upper Confidence Bound (UCB), Thompson sampling, among others. For a deeper introduction to these methods and their deployment on AWS, see our earlier post [Dynamic A/B testing for machine learning models with Amazon SageMaker MLOps Projects](https://aws.amazon.com/blogs/machine-learning/dynamic-a-b-testing-for-machine-learning-models-with-amazon-sagemaker-mlops-projects/).\n\nIn Amazon Payments, we chose UCB, a strategy that selects the arm with the highest estimated reward plus an uncertainty bonus, naturally balancing exploitation and exploration. Its deterministic selection rule gives an auditable, reproducible decision for every impression, which makes sure every serving decision can be explained and reproduced if needed. However, a standard bandit learns one best arm for the entire audience. Personalization requires conditioning on who the visitor is, and that is what contextual bandits provide.\n\n### Adding the context using contextual bandits\n\nInstead of asking “which content is the best overall?,” a contextual bandit conditions its decision on signals about the visit. Segmented bandits run a separate instance per hand-defined group, but the groups are arbitrary and each needs its own traffic. Contextual bandits condition directly on a feature vector (a numeric representation of attributes), so a pattern learned in one context transfers to every similar visit without separate per-group traffic. Our production system represents each customer as a context vector of behavioral signals (payment behavior, transaction mix, and similar features) in place of a fixed segment. The entity ID (an opaque key such as `entity_id`) is used only to route the learned recommendation back to the right visitor. It is never a model input.\n\n### LinUCB: Learning at the feature level\n\nFor a contextual approach, we selected Linear UCB (LinUCB), introduced by [Li et al. (2010)](https://arxiv.org/pdf/1003.0146). While newer bandit algorithms exist, LinUCB remains a battle-tested method that is computationally efficient, auditable (deterministic arm selection), and naturally handles a large arm space with limited warm-start data. Its key assumption is that the expected reward for an arm is a linear function of the context vector, which lets it generalize to visitors it has not seen before.\n\nEach arm keeps two running tallies updated on every impression:\n\n- **`b`, the reward ledger:** which visitor signals led to conversions (b += reward · x).\n- **`A`, the experience ledger:** which visitors the arm has seen (A += x·xᵀ, starting from identity).\n\nDividing reward by experience gives the estimate (θ = A⁻¹·b). The experience ledger also shrinks the exploration bonus as evidence grows. The arm’s score is:\n\nThe bonus is context-dependent. It’s large for a kind of visitor the arm has rarely seen, and small for one it has seen often. A plain bandit explores at a single global rate, while LinUCB adjusts how much it explores for every visitor it scores.\n\n## Optimizing an entire funnel\n\nThe customer journey in our use case comprises three steps: application start, submission, and approval. The objective is therefore a multi-stage outcome, not a single metric. These stages do not move in unison. Content optimized for starts tends to attract a broad audience, yet approval depends on whether the offer genuinely suits the applicant. Optimizing one stage in isolation can degrade another. We refer to this as the seesaw problem. Conversely, optimizing solely for approvals starves the model of signal, because approvals are rare and delayed.\n\nOur approach optimizes the entire funnel simultaneously by running one LinUCB model per stage (start, submit, approve) and combining their UCB scores through a linear combination:\n\nThe stage weights can be assigned based on business priorities or learned by a separate calibration step. We used approximately equal weights. In practice, one could weight approvals more heavily once the model is warmed or use a Pareto frontier if the trade-off is genuinely contested.\n\nHere’s the core of the model in Python. Each LinUCBDisjoint maintains independent parameters per arm, and MultiObjectiveLinUCB composes three of them:\n\nWe built [a quick start Jupyter notebook](https://github.com/aws-samples/sample-multi-objective-contextual-bandit-with-sagemaker/blob/main/notebooks/walkthrough.ipynb) to test this approach on Amazon SageMaker AI. [The repository](https://github.com/aws-samples/sample-multi-objective-contextual-bandit-with-sagemaker) also includes a command-line demo and unit tests.\n\n**Tuning the alpha.** The α parameter controls the exploration-exploitation balance. Higher values encourage exploration of under-tested arms, and lower values favor exploitation. A value of α = 1.0 is a reasonable default (Li et al., 2010). In practice, start higher when the arm space is large and history is limited, then reduce α as evidence accumulates. A typical range is 0.1 to 2.0. Values that are too high waste traffic on weak arms, and values that are too low risk locking onto a suboptimal arm. Because α is a single scalar, a grid search is straightforward.\n\n**Handling delayed feedback.** In our setting, approval decisions lag by days, so the reward for the final funnel stage isn’t observed until well after the impression. We address this with an attribution window. Starts and submissions update the model immediately, while approval outcomes are held back until the subsequent batch cycle, avoiding downward bias from pending applications. This aligns naturally with our weekly processing cadence.\n\n## Publishing safe content at scale\n\nA bandit is only as good as the pool of arms. Building a rich arm pool responsibly is the next problem to solve at scale while maintaining content oversight.\n\nThe key idea is to compose many variations from a small set of reviewed building blocks. Our content was assembled from two kinds of parts: *industry-themed images* and *benefit-focused taglines*. The arm space is the Cartesian product of these parts, so each arm is one (image, tagline) pairing and a modest number of building blocks produces a large pool of distinct page variations.\n\nThat structure also makes the approach scalable without sacrificing oversight, through a few layers of control.\n\n**Vetting the parts, not the combinations.** Some content elements are fixed by stakeholder requirements. Rather than reviewing every possible combination, which would be impractical as the space grows, we vet each individual building block up front. Because the number of building blocks remains small and reviewable while the combinations grow rapidly, vetting the parts once provides assurance over the entire combinatorial space.\n\n**Design systems for consistency.** As described in our earlier post, anchoring content to a design system (approved colors, layouts, components, and patterns) keeps variations visually consistent and on-brand by construction, not by chance.\n\n**Expanding the pool with generative AI.** In our initial deployment, building blocks were curated with detailed human oversight, even when production was assisted by generative AI tools. We see clear potential to increase the parts pool (text, imagery, layouts) substantially, which would widen the bandit’s selection space. As this pool grows, the bandit and the generative pipeline form a virtuous cycle. Generative AI widens the candidate set, and the bandit identifies which combinations yield the strongest outcomes per visitor.\n\n## Deploying on AWS\n\nAmazon SageMaker AI is an AI model development service that supports the full model lifecycle. It’s the primary AWS service behind our end-to-end pipeline, providing model training, batch processing, version control, and monitoring in a unified tool. We adopted a batch architecture for two reasons: the offline nature of the selection problem (feedback accumulates over days), and the observation that visitor selection behavior does not shift rapidly enough to require real-time model updates. A scheduled SageMaker AI Processing job reads the prior period’s feedback, updates the model, and writes fresh recommendations for the next period.\n\n### Architecture overview\n\nWe run a weekly SageMaker AI batch job that reads customer outcomes from [Amazon Simple Storage Service (Amazon S3)](https://aws.amazon.com/s3/), updates the bandit model, and publishes per-customer recommendations to a low-latency key-value store. Starting with a weekly cadence provides a conservative baseline. Increasing frequency is straightforward once the model demonstrates stable lift. Approval feedback, which can lag by days, is incorporated into subsequent batches.\n\n**1. Data collection.** Customer impressions and outcomes (starts, submissions, approvals) are logged at end-of-session and flow into Amazon S3 as observations.\n\n**2. Model update and inference**. An Amazon SageMaker AI Processing job runs the Python entry point. It loads the most recent model state from Amazon S3, discovering the latest dated model prefix automatically so that nothing needs reconfiguring between runs. It then splits the data into feedback and inference rows, updates the model on the feedback (incremental updates), and scores every prospect to select an arm. We use a Processing job because our workload is a single, self-contained step that both updates and scores, with custom S3 I/O. Neither a Training job nor Batch Transform maps cleanly to this pattern.\n\n**3. Persistence and output.** The job writes the updated model state back to Amazon S3 (in a new dated path, which gives a built-in version history and straightforward rollback), alongside the observations and the arm content. The per-customer recommendations (each customer mapped to their selected arm) are published to a low-latency key-value store (such as [Amazon DynamoDB](https://aws.amazon.com/dynamodb/)) that fronts live traffic.\n\n**Warm-starting the model.** We warm-started each model from a period of randomized content assignment, which provides unbiased data that lets the bandit initialize its beliefs without selection bias and spend its exploration budget more efficiently from day one.\n\n**Scaling inference efficiently.** Scoring a large prospect population is compute-intensive. Two optimizations kept it within budget: precomputing the arm matrix inversions once (they do not change within a batch), and splitting the prospect set into chunks scored in parallel using Python’s `multiprocessing.Pool`.\n\n**How the page is served.** The serving path is straightforward because the heavy lifting has already happened in the batch job. The personalized page is delivered by a low-latency key-value store (such as Amazon DynamoDB) that holds each customer’s precomputed recommendation. When a customer arrives, the page performs a single lookup by entity ID and renders the precomputed arm’s content. There is no real-time model inference.\n\nFor use cases that require real-time scoring, where context is only known at request time or content must adapt within a session, Amazon SageMaker AI real-time inference endpoints are the natural alternative, hosting the model behind a low-latency API.\n\nThe design has one more robustness property worth calling out. If no recommendation exists for a visitor, the page falls back to the default static experience, so no customer is worse off than baseline. This fallback bounds the downside and makes incremental rollout safer.\n\n## Lessons learned\n\nDeploying this system in production taught us several things that generalize beyond our specific use case.\n\n### Model design\n\n*When populations differ a lot, train separately.* If your audience splits into groups with significantly different features or conversion rates, a single shared model lets the larger or noisier group dominate what gets learned. Training one model per group, giving each its own feature set and arm space, lets each audience’s signal come through cleanly.\n\n*Optimize the funnel, not a proxy.* The single most important design decision was going multi-objective. A single-stage bandit would have let us “win” on an upstream metric while degrading performance on the downstream metric that ultimately determined business value. We validated this empirically before deployment. Policies optimized for a single funnel stage consistently produced at least one negative directional lift elsewhere. The multi-objective formulation was the only approach that maintained non-negative estimates across all three stages simultaneously. If your business goal is several steps removed from the first click, encode all the steps in the reward.\n\n### Content strategy\n\nThe selection algorithm addresses only half of the challenge. The quality of the content pool is equally important. In our A/B test, one customer population saw directionally positive lifts across all three funnel metrics, currently a high single-digit percentage relative lift on the final stage. For the second population, the model explored most of its arm pool yet found no combination that beat the static page. The lifts were negative, and the approval regression was statistically significant. This was not a model failure. Thorough exploration can confirm that a content pool contains no winner.\n\nIf a bandit explores the arms broadly and still cannot beat the control (the static experience, with no bandit), then the arms need revisiting instead of the algorithm. This is where generative content pipelines, like the one in our [earlier post](https://aws.amazon.com/blogs/machine-learning/reinvent-personalization-with-generative-ai-on-amazon-bedrock-using-task-decomposition-for-agentic-workflows/), become essential. They expand and refresh the pool the bandit draws from, so the model has a better chance of finding a winning combination.\n\nConsider making the baseline an arm. Including the existing default experience as one of the arms means that if no personalized arm beats the default for a given customer, UCB naturally gravitates toward serving it, bounding how far the system can regress.\n\n### Evaluation\n\nOffline evaluation and A/B testing answer different questions. Offline replay (policy compared to random on held-out data) tells you “does the model beat random assignment?”, serving as an initial validation. But only a conventional A/B test (bandit-personalized compared to static baseline) determines whether personalization outperforms your existing experience. We used both. Offline estimates gave confidence to launch, and the A/B test delivered the definitive verdict, including the finding that content pool quality was the binding constraint on one population.\n\nThe weights tell you what is driving the outcome. LinUCB’s learned θ vectors are directly interpretable. A large positive weight on a feature means customers strong on that signal are more likely to convert on that arm. After a few weeks you can ask “what signals differentiated responders to this image?” and get credible insights. This yields built-in interpretability without additional tooling, and provides a direct input to content strategy.\n\n### Operations\n\nBatch is a legitimate deployment pattern. You do not need a real-time online-learning application to run a bandit in production. A weekly Processing job reading feedback, updating an exact-update model, and writing recommendations to S3 was auditable, resource-efficient, and a natural fit for delayed feedback.\n\n## Conclusions\n\nGenerative AI removed the constraint on producing personalized content. The remaining challenge is deciding which content to show whom and learning the answer from real customer feedback. Multi-armed bandits are purpose-built for this. With a contextual, multi-objective formulation, you can personalize on individual behavior while optimizing an entire conversion funnel.\n\nThe system is straightforward to operate on AWS. A scheduled SageMaker AI Processing job plus Amazon S3 for versioned model state gives you incremental learning, an audit trail, and a safe fallback without requiring a specialized real-time tool.\n\nIf generative AI is already filling your content pool, a contextual bandit is the natural optimization layer on top, identifying for each individual which content variant yields the strongest outcome. The two techniques in this post series are complementary.\n\nTo get started, open the [Amazon SageMaker AI console](https://console.aws.amazon.com/sagemaker/), explore the [SageMaker AI developer guide](https://docs.aws.amazon.com/sagemaker/), clone the [code sample](https://github.com/aws-samples/sample-multi-objective-contextual-bandit-with-sagemaker) and open its walkthrough notebook to run the bandit yourself, and review our earlier post on [generative AI personalization on Amazon Bedrock](https://aws.amazon.com/blogs/machine-learning/reinvent-personalization-with-generative-ai-on-amazon-bedrock-using-task-decomposition-for-agentic-workflows/).", "url": "https://wpnews.pro/news/uplifting-conversion-across-the-acquisition-funnel-with-personalization-using-on", "canonical_source": "https://aws.amazon.com/blogs/machine-learning/uplifting-conversion-across-the-acquisition-funnel-with-personalization-using-contextual-bandits-on-aws/", "published_at": "2026-10-01 16:51:04+00:00", "updated_at": "2026-10-01 17:15:19.207263+00:00", "lang": "en", "topics": ["machine-learning", "ai-products", "ai-infrastructure", "artificial-intelligence"], "entities": ["Amazon Payments", "Amazon SageMaker AI", "Amazon Bedrock", "Amazon Web Services", "Upper Confidence Bound"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/uplifting-conversion-across-the-acquisition-funnel-with-personalization-using-on", "markdown": "https://wpnews.pro/news/uplifting-conversion-across-the-acquisition-funnel-with-personalization-using-on.md", "text": "https://wpnews.pro/news/uplifting-conversion-across-the-acquisition-funnel-with-personalization-using-on.txt", "jsonld": "https://wpnews.pro/news/uplifting-conversion-across-the-acquisition-funnel-with-personalization-using-on.jsonld"}}