{"slug": "zalando-let-ai-nudge-its-prices-sales-didnt-flinch", "title": "Zalando Let AI Nudge Its Prices. Sales Didn’t Flinch", "summary": "Zalando tested a daily machine-learning pricing system across 23 A/B tests in 12 European markets and found it improved profit during sales without reducing revenue or units sold after returns, according to a company-authored preprint submitted June 11, 2026. The Berlin-founded fashion platform replaced a weekly Transformer forecaster with a LightGBM demand model feeding a linear-programming optimizer that balances net merchandise value against long-term profit, covering roughly 600,000 products. Zalando ran the tests without showing different customers different prices, addressing a mismatch between sales events that average 6.2 days and a pricing workflow that previously took hours per decision.", "body_md": "8 min read\n\n**Zalando replaced a slow pricing workflow with daily machine-learning forecasts, then tested the system across 12 European markets without showing different customers different prices.**\n\nA sales event lasts an average of 6.2 days at [Zalando](https://corporate.zalando.com/en/about-us). Its old pricing system worked by the week.\n\nThat mismatch mattered because sales campaigns occupy about 35% of Zalando’s year. Demand moves sharply during an end-of-season sale or Black Friday campaign, discounts change more often, and pricing teams may need to reconsider hundreds of thousands of products before the event ends. The existing workflow combined model outputs, manual analysis and heuristic rules. A pricing decision could take hours.\n\nThe Berlin-founded fashion platform built a new system around daily machine learning demand forecasting and a multi-objective optimizer. The question was concrete. Could Zalando algorithmic pricing improve profit during a sale without reducing revenue or the number of products sold after returns?\n\nA [company-authored preprint submitted on June 11, 2026](https://arxiv.org/abs/2606.13741) describes how Zalando answered it through 23 A/B tests in 12 European markets. The result is less a story about automated discounting than about how a European retailer tested a consequential AI system across markets that did not behave alike.\n\n## Weekly pricing could not keep up with a six-day sale\n\nZalando manages prices for roughly 600,000 products at any time. During ordinary trading periods, its existing system forecast demand and costs for possible discounts, then chose a markdown between zero and 70%. The production forecaster used a Transformer model with weekly resolution and weekly updates.\n\nSales campaigns broke the cadence. Discounts during these events averaged five percentage points deeper than on other days. Upload frequency and the variation between discounts both rose. A weekly forecast could miss a demand jump that appeared and disappeared inside the same campaign.\n\nThe workflow also treated financial outcomes indirectly. Pricing analysts combined algorithmic runs, manual adjustments and rules designed around target discount levels. Revenue and long-term profit influenced those decisions, but they were not explicit objectives inside one optimization problem.\n\nZalando wanted faster decisions and a clearer trade-off. The new forecast then optimize architecture separated the work into two stages. A LightGBM model forecast demand for combinations of products, markets, dates and discount levels. A linear-programming optimizer then selected article discounts using a weighted objective that balanced net merchandise value with long-term profit.\n\n## A smaller model was better suited to the decision\n\nThe team compared gradient-boosted trees, two neural-network approaches, a naive model and a daily version derived from Zalando’s weekly production forecaster. The comparison covered 212 days between May and December 2024, including every major type of sales event.\n\nThe gradient-boosted model performed best on Zalando’s internal demand and gross-merchandise-value error measures. The neural models posted similar results on standard error measures, but the downstream decision was the point. A forecast that scores well in isolation can still lead an optimizer toward poor prices.\n\nZalando simulated that distinction by comparing mean-squared-error and Tweedie loss functions. The mean-squared-error model produced more optimistic forecasts, then delivered materially less than it had projected when its outputs passed through the optimizer. The Tweedie model forecast more conservatively and produced outcomes closer to those forecasts.\n\nThis is the first transferable lesson from the machine learning demand forecasting work. Model selection followed the decision the forecast would support. Zalando did not choose a forecaster solely because it minimized a familiar statistical error. It checked what happened after optimization converted predictions into discount recommendations.\n\nThe resulting system cut pricing-decision time from hours to minutes. It remained a decision-support tool. Pricing managers retained oversight and could inspect alternative trade-offs rather than handing financial accountability to a fully autonomous process.\n\n## Products, not shoppers, were randomized\n\nPricing creates an experimental-design problem that a normal customer-level split cannot solve cleanly. If two shoppers see different prices for the same product at the same time, the test creates price discrimination and a poor customer experience. Zalando randomized articles instead.\n\nZalando divided products entering each ecommerce pricing experiment into two equally sized groups. Control articles kept discounts produced by the previous system with manual modifications. Treatment articles received discounts recommended by the new algorithm.\n\nSimple article-level randomization could still produce imbalance or interference. Similar products may substitute for one another, and one group could end up with a different historical contribution to revenue and profit. Zalando used clustered randomization to reduce substitution between similar products while balancing product characteristics and pre-intervention financial variables.\n\nThe team treated adoption as a non-inferiority problem first. The new system needed to perform at least as well as the old workflow on net merchandise value and profitability. Matching the old commercial outcome would still justify adoption because the system reduced operational complexity and compressed the decision cycle.\n\n## Twenty-three tests covered twelve European markets\n\nZalando ran 23 tests during sales campaigns in 2023 and 2024. They covered Belgium, Denmark, Finland, France, Germany, Italy, the Netherlands, Norway, Poland, Spain, Sweden and Switzerland. Experiments lasted four to ten days, with an average of seven days. Individual waves included about 200,000 to more than 1.8 million articles, and the program covered 6.2 million article observations in total.\n\nThis European retail AI test could not assume that one country represented the rest. Assortment, demand, sales-event timing and data volume differ by market. Zalando first used difference-in-differences analysis within the experiments to account for temporal trends and market-level factors. It then combined the waves through a Bayesian hierarchical model designed to preserve variation across markets and periods.\n\nAcross the portfolio, the company reported a 6% increase in profit contribution, with a 95% credible interval from 0.79% to 11.03%. Net merchandise value rose an estimated 2.23%, but its interval ran from minus 0.83% to 5.33%. Sales after returns moved 0.56%, with an interval from minus 3.10% to 4.26%.\n\nThe available evidence therefore supports a profit improvement under the company’s model while revenue and sales remained statistically equivalent to the previous process. It does not support a general claim that the algorithm increased every commercial outcome.\n\nGermany produced larger reported effects. Sales after returns increased 16.7%, net merchandise value increased 17.9%, and profit contribution increased 16.7%, all described as statistically significant. Zalando suggests that scale and richer data may explain the stronger outcome. That explanation is an inference, not a mechanism isolated by the test.\n\n## The market average did not erase market differences\n\nThe European lens changes what operators should take from the result. A company expanding one AI decision system across multiple countries needs an estimate of the overall effect and a view of where that estimate varies. Pooling every observation into one average can hide market differences. Reading each country independently can leave smaller markets too underpowered to guide a decision.\n\nThe hierarchical analysis offered a middle route. It allowed the company to estimate a portfolio effect while accounting for heterogeneity between test waves. That matters when the operating decision is a shared system rollout rather than a collection of unrelated local launches.\n\nThe design still leaves important questions open. The paper provides detailed results only for Germany and does not show every country’s posterior estimate. It does not disclose absolute profit, revenue or discount levels. Article-level assignment reduces customer-facing price inconsistency, but products can still substitute for one another. The model also omits explicit cross-price elasticity and demand uncertainty.\n\n## A profit lift was enough to change the workflow\n\nAfter the test program, Zalando deployed the algorithm to production. The company says it now handles the majority of algorithmic pricing decisions for sales campaigns across its markets. Pricing teams moved toward strategic oversight while the system performed routine discount calculations.\n\nThe experiment offers a practical sequence for other European operators building AI into consequential business decisions. Begin with an explicit financial objective and a guardrail that protects the existing business. Evaluate forecasts through the decisions they produce. Randomize at a level that avoids unfair treatment. Preserve country-level variation when combining results. Keep human accountability where the trade-off remains a business judgment.\n\nZalando algorithmic pricing did not remove the conflict between revenue, inventory and profit. It turned that conflict into an objective the system could display and an experiment the company could inspect. That distinction explains why the test changed the operating workflow rather than ending as another model-performance report.\n\n## Sources and limits\n\n- [Zalando’s high-frequency pricing preprint](https://arxiv.org/abs/2606.13741) , submitted June 11, 2026\n- [Zalando’s official company overview](https://corporate.zalando.com/en/about-us)\n- [Lead author Stefan Birr’s LinkedIn profile](https://de.linkedin.com/in/stefan-birr-30471a160)\n\nZalando employees produced the experiment and paper, and one co-author was affiliated with Databricks when it was written. The results have not been independently replicated. The paper reports modeled relative effects and intervals but withholds absolute commercial outcomes, full market-level results and the exact discounts assigned to treatment and control articles.", "url": "https://wpnews.pro/news/zalando-let-ai-nudge-its-prices-sales-didnt-flinch", "canonical_source": "https://industrycontents.com/zalando-algorithmic-pricing/", "published_at": "2026-09-28 08:00:00+00:00", "updated_at": "2026-09-28 08:18:26.472501+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence", "ai-research"], "entities": ["Zalando", "LightGBM", "Transformer", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/zalando-let-ai-nudge-its-prices-sales-didnt-flinch", "markdown": "https://wpnews.pro/news/zalando-let-ai-nudge-its-prices-sales-didnt-flinch.md", "text": "https://wpnews.pro/news/zalando-let-ai-nudge-its-prices-sales-didnt-flinch.txt", "jsonld": "https://wpnews.pro/news/zalando-let-ai-nudge-its-prices-sales-didnt-flinch.jsonld"}}