6 min read
o2 could not profile most of its visitors, so it personalised from what each session revealed.
Picture a shop assistant who has never met you and still has to guess what you want. That was the position of o2’s online shop, where most commercially valuable visitors were anonymous. Without a login or durable customer history, Telefónica Deutschland could not use the usual customer relationship management (CRM) profile to decide what to show them. Its older recommendation system fell back on standard offers and rules that staff maintained by hand.
The German telecom company tested whether AI product recommendations could make more of the weak signals available inside a browsing session. According to a case study published by implementation partner e-dialog, personalised hero banners, the large panels at the top of the home page, produced a 17% click-through rate and a 275% conversion-rate increase in an A/B test.
The operating lesson matters more than the headline number. Telefónica worked with the session in front of it. It built an event-driven pipeline that refreshed the recommendation whenever the visitor’s behaviour changed.
Static rules could not follow a changing session #
Telefónica Deutschland, headquartered in Munich and best known for the o2 brand, sells mobile and fixed-line services. The growth problem was concentrated on its website, where the majority of commercially valuable traffic was not logged in, according to e-dialog.
That left the old system with two options. It could show a default offer to everyone or apply manually maintained rules. Neither responded well when a shopper changed direction during a visit, such as moving from iPhone pages to Samsung devices. Retailers face a similar shortage of durable data when AI shopping agents arrive knowing intent but not the shopper.
The team gives no formal hypothesis, but the documented design implies one. Current-session behaviour could be enough for anonymous visitor personalization if the system could ingest those events and update the product ranking before the next page loaded.
A small accessory test proved the pipeline #
Telefónica and e-dialog started with accessory recommendations on product pages. Before the change, specialists prepared rankings in spreadsheets, converted them to JSON, a standard data format, and uploaded them. The proposed replacement joined Google Analytics behavioural events with current product-feed data and passed both into Google Cloud’s Vertex AI Retail API.
The model returned a ranked JSON output to the product page. e-dialog says this first use case removed the manual ranking work and increased online-store conversions by 66%, with return on investment above 550%. The test duration, sample size, traffic split and confidence interval behind that pilot result all go undisclosed.
The pilot mattered because it tested the data path before the team used the system in its highest-reach placement. A growth operator can copy that move. Validate one narrow recommendation surface, including its feed freshness and event timing, before asking the same pipeline to control a prominent part of the journey.
The hero banner reacted to live intent #
The second phase moved Vertex AI recommendations into the hero banners on o2’s home page. Instead of reading a stored profile, the system used clicks and browsing behaviour from the active session. When the pattern changed, the model recalculated the likely next product and updated the banner.
e-dialog describes the architecture as event driven and says it was defined as infrastructure as code, meaning the cloud setup is written as code files and deployed through an automated pipeline. That made the setup auditable and reusable across other brands or use cases. Product-catalogue data and user events were continuously refreshed rather than assembled for each campaign.
This went well beyond a creative swap. The intervention combined a live event stream, a current product catalogue, a ranking model and a delivery layer fast enough to act on the next page. Remove any one of those pieces and real-time personalization becomes a delayed campaign. AMap’s homepage recommendation test showed a similar pattern, where a faster model lifted clicks on a single placement.
The reported lift needs a clearer baseline #
e-dialog says an A/B test validated the hero experience after millions of recommendations. The personalised banners recorded a 17% click-through rate, a 275% conversion rate uplift and a reported 939% return on investment. The article also refers to millions of calculated recommendations in 30 days, although its text gives no exact total.
Steven Burkhardt, Telefónica Deutschland’s Head of Digital Analytics and Performance Management, is quoted in the case and has publicly shared that his team tested how AI-generated product recommendations affected purchase behaviour. That corroborates the project and its experimental intent.
Several important evidence gaps remain. Neither source gives enough detail on the control experience, and neither shares the traffic allocation, absolute conversion rates, statistical confidence, customer segment or test dates. The 275% figure is therefore a relative, partner-published claim. A reader cannot tell whether conversion moved from a very small baseline or how much incremental profit survived implementation costs.
Copy the sequence before copying the model #
A comparable team can test the method without rebuilding Telefónica’s cloud stack. Start with one decision point where anonymous visitors reveal fresh intent. Define the events that represent a meaningful change, check how quickly they arrive, and make the control explicit. Keep the old default or rule-based recommendation running for a randomly assigned holdout, a group that never sees the new system.
Track four layers separately. These are event freshness, recommendation delivery, interaction with the placement and the final commercial conversion. Add margin or contribution profit before calling the test a business win. A recommendation can raise clicks while sending customers toward lower-value products, and a conversion lift can fail to cover platform and engineering costs.
The most useful follow-up would compare the real-time model with a strong rules-based system and go beyond a generic banner. A second holdout could delay the same recommendation by one page to measure whether speed itself creates value. That separates the contribution of the ranking model from the value of a fresher data pipeline.
Telefónica’s reported result suggests that an anonymous session may contain enough intent to improve a high-value placement. The hard part sat in making weak signals usable quickly and then connecting the decision to an experiment that could measure a business outcome.
Sources #
Compare the evidence and operating lessons from more AI experiments in our Growth Signal Index, or browse benchmark tooling in the Industry Contents Lab.
Spotted an error? See our corrections policy.