{"slug": "allegros-ai-had-a-good-idea-it-just-picked-the-wrong-spot", "title": "Allegro’s AI Had a Good Idea. It Just Picked the Wrong Spot", "summary": "Allegro's AlleCompanion machine-learning retrieval system lifted attributed gross merchandise value by as much as 21.25% in the in-cart placement but left GMV flat in the pre-cart step immediately before it, according to a company-authored paper released September 4, 2026. The Polish online marketplace, which serves more than 20 million active users a month, ran five two-week online experiments covering sponsored and organic product-page modules, two pre-cart strategies and an in-cart module, routing all platform traffic and comparing AlleCompanion against the existing production recommendation mix. AlleCompanion combined a category-conditioned two-tower recommender with behavioural filters, expert compatibility rules, human annotation, statistical mining and an LLM-assisted category map to separate complementary purchases from substitutes.", "body_md": "19 min read\n\n**Allegro tested one machine-learning retrieval system across five shopping placements. It lifted attributed GMV (gross merchandise value, the value of goods sold) by as much as 21.25% in the cart, yet left GMV flat in the pre-cart step right before it.**\n\nYou add a camera to your basket. What should the shop suggest next? A lens, a tripod or a compatible case would all help. Another camera body only helps if you are still comparing cameras.\n\nAt [Allegro](https://about.allegro.eu/), a Polish online marketplace, that distinction became a data problem. Purchase logs record what people bought together and stay silent on why. Two flavours of dog food may share a transaction because they are alternatives. A phone and a case may share one because the case completes the phone. A recommendation model trained on both patterns learns co-purchase frequency and misses complementarity.\n\nThe company built AlleCompanion to separate those signals. This Allegro recommendation system combined a category-conditioned two-tower recommender with behavioural filters, expert compatibility rules, human annotation, statistical mining and an LLM-assisted category map. A two-tower model converts each item into a set of numbers and retrieves the closest matches. An LLM, or large language model, is the type of AI behind modern chatbots.\n\nThe test that mattered came after the offline model work. Would the same retrieval system create commercial value wherever Allegro placed it, or would shopper intent change the answer from one part of the journey to the next?\n\nA [company-authored paper released on September 4, 2026](https://arxiv.org/html/2609.05063v1) reports five online experiments. They covered sponsored (paid) and organic product-page modules, two pre-cart strategies and an in-cart module. Each ran for two weeks, routed all platform traffic and compared AlleCompanion with the existing production recommendation mix.\n\n## Co-purchase data confused complements with alternatives\n\nAllegro serves more than 20 million active users a month. Its catalogue and transaction volume provide a large behavioural dataset, yet more data left the ambiguity in place.\n\nThe team defined co-purchase sessions and turned them into ordered item pairs, which preserves the direction of a relationship. A smartphone can create demand for a case, while a case rarely creates demand for a smartphone. The team tuned the session window against a holdout set (a slice of data kept aside for testing) because short windows produced many near-identical items and long windows pulled in unrelated purchases.\n\nThey also excluded buyers above the 99th percentile for transaction volume. Those heavy buyers could distort the distribution and rarely reflect ordinary shopping intent.\n\nAn internal annotation exercise exposed the remaining noise. Two annotators classified each of 400 sampled item pairs as complementary, substitutable or unrelated. Raising the minimum number of times a pair had appeared cut unrelated examples from 44% to 22%, while substitutes rose from 28% to 43%. Frequency filtered out accidents and strengthened a different kind of error.\n\nA category constraint changed the mix more directly. Requiring two items to share a department but sit in different categories moved the sample from 36% complementary, 16% substitutable and 48% unrelated to 61% complementary, 4% substitutable and 35% unrelated.\n\nThat improved the mix and still left a gap. Category labels could say that phone cases complement phones, yet they could not prove that a particular case fits a particular model. The team needed category-level intent and item-level compatibility at the same time.\n\n## Allegro separated category policy from item retrieval\n\nAlleCompanion divided the problem into two layers. The retrieval model learned item relationships. A separate Complementary Categories Mapping, called ComCat, told the model which category relationships suited a placement.\n\nComCat merged three sources. Automated heuristics mined directional category relationships from purchases. Human annotators reviewed pairs drawn from traffic and an LLM-assisted workflow, including low-traffic categories. Business experts supplied rules for technical compatibility that purchase behaviour might miss.\n\nThe sources followed a priority order. Human annotations came first, expert rules next and automated heuristics served as the broadest fallback. Allegro could also add a same-category relation when a placement needed alternatives as well as complements.\n\nSeparating the layers let Allegro change the category policy and skip retraining the underlying retrieval model. For AI product recommendations spread across several surfaces, product intent becomes a configurable layer instead of a property frozen into the model weights.\n\nThe model architecture enforced the category signal during retrieval. A conventional two-tower baseline mostly returned similar items. A post-processing alternative retrieved a broad candidate set and filtered it afterward, yet even after oversampling candidates by 15 times, the median final list held only seven items. Allegro’s Category Adapter guided candidate generation inside the embedding space, the map of numbers where similar items sit close together.\n\nThe researchers trained comparable variants on 90 days of data and evaluated them on a later seven-day holdout period. They kept hyperparameters, the tuning settings, identical across model configurations. AlleCompanion improved offline retrieval measures. In the paper’s phone example, it returned cases compatible with the exact device and skipped other phones or cases designed for a different model.\n\n## Strictly cleaner training data performed worse\n\nAllegro then tested whether cleaner data would improve the model. One dataset filtered transaction pairs through expert compatibility rules. Another generated synthetic pairs from those rules. A third used synthetic data for pretraining and transaction data for fine-tuning.\n\n**Allegro tested one machine-learning retrieval system across five shopping placements. It lifted attributed GMV (gross merchandise value, the value of goods sold) by as much as 21.25% in the cart, yet left GMV flat in the pre-cart step right before it.**\n\nYou add a camera to your basket. What should the shop suggest next? A lens, a tripod or a compatible case would all help. Another camera body only helps if you are still comparing cameras.\n\nAt [Allegro](https://about.allegro.eu/), a Polish online marketplace, that distinction became a data problem. Purchase logs record what people bought together and stay silent on why. Two flavours of dog food may share a transaction because they are alternatives. A phone and a case may share one because the case completes the phone. A recommendation model trained on both patterns learns co-purchase frequency and misses complementarity.\n\nThe company built AlleCompanion to separate those signals. This Allegro recommendation system combined a category-conditioned two-tower recommender with behavioural filters, expert compatibility rules, human annotation, statistical mining and an LLM-assisted category map. A two-tower model converts each item into a set of numbers and retrieves the closest matches. An LLM, or large language model, is the type of AI behind modern chatbots.\n\nThe test that mattered came after the offline model work. Would the same retrieval system create commercial value wherever Allegro placed it, or would shopper intent change the answer from one part of the journey to the next?\n\nA [company-authored paper released on September 4, 2026](https://arxiv.org/html/2609.05063v1) reports five online experiments. They covered sponsored (paid) and organic product-page modules, two pre-cart strategies and an in-cart module. Each ran for two weeks, routed all platform traffic and compared AlleCompanion with the existing production recommendation mix.\n\n## Co-purchase data confused complements with alternatives\n\nAllegro serves more than 20 million active users a month. Its catalogue and transaction volume provide a large behavioural dataset, yet more data left the ambiguity in place.\n\nThe team defined co-purchase sessions and turned them into ordered item pairs, which preserves the direction of a relationship. A smartphone can create demand for a case, while a case rarely creates demand for a smartphone. The team tuned the session window against a holdout set (a slice of data kept aside for testing) because short windows produced many near-identical items and long windows pulled in unrelated purchases.\n\nThey also excluded buyers above the 99th percentile for transaction volume. Those heavy buyers could distort the distribution and rarely reflect ordinary shopping intent.\n\nAn internal annotation exercise exposed the remaining noise. Two annotators classified each of 400 sampled item pairs as complementary, substitutable or unrelated. Raising the minimum number of times a pair had appeared cut unrelated examples from 44% to 22%, while substitutes rose from 28% to 43%. Frequency filtered out accidents and strengthened a different kind of error.\n\nA category constraint changed the mix more directly. Requiring two items to share a department but sit in different categories moved the sample from 36% complementary, 16% substitutable and 48% unrelated to 61% complementary, 4% substitutable and 35% unrelated.\n\nThat improved the mix and still left a gap. Category labels could say that phone cases complement phones, yet they could not prove that a particular case fits a particular model. The team needed category-level intent and item-level compatibility at the same time.\n\n## Allegro separated category policy from item retrieval\n\nAlleCompanion divided the problem into two layers. The retrieval model learned item relationships. A separate Complementary Categories Mapping, called ComCat, told the model which category relationships suited a placement.\n\nComCat merged three sources. Automated heuristics mined directional category relationships from purchases. Human annotators reviewed pairs drawn from traffic and an LLM-assisted workflow, including low-traffic categories. Business experts supplied rules for technical compatibility that purchase behaviour might miss.\n\nThe sources followed a priority order. Human annotations came first, expert rules next and automated heuristics served as the broadest fallback. Allegro could also add a same-category relation when a placement needed alternatives as well as complements.\n\nSeparating the layers let Allegro change the category policy and skip retraining the underlying retrieval model. For AI product recommendations spread across several surfaces, product intent becomes a configurable layer instead of a property frozen into the model weights.\n\nThe model architecture enforced the category signal during retrieval. A conventional two-tower baseline mostly returned similar items. A post-processing alternative retrieved a broad candidate set and filtered it afterward, yet even after oversampling candidates by 15 times, the median final list held only seven items. Allegro’s Category Adapter guided candidate generation inside the embedding space, the map of numbers where similar items sit close together.\n\nThe researchers trained comparable variants on 90 days of data and evaluated them on a later seven-day holdout period. They kept hyperparameters, the tuning settings, identical across model configurations. AlleCompanion improved offline retrieval measures. In the paper’s phone example, it returned cases compatible with the exact device and skipped other phones or cases designed for a different model.\n\n## Strictly cleaner training data performed worse\n\nAllegro then tested whether cleaner data would improve the model. One dataset filtered transaction pairs through expert compatibility rules. Another generated synthetic pairs from those rules. A third used synthetic data for pretraining and transaction data for fine-tuning.\n\nThe expert-filtered dataset removed 77% of the available training pairs. It improved attribute consistency and weakened standard relevance measures. Models trained only on synthetic expert-rule pairs also performed poorly. The strongest compromise used synthetic pretraining followed by fine-tuning on behavioural data.\n\nLabels that look cleaner to experts can remove the behavioural variation a model needs in production. Allegro resolved that tension by moving precision into the controllable category layer and leaving the retrieval model exposed to much of the real distribution.\n\n## Five tests put the same system in five contexts\n\nThe ecommerce recommendation experiment compared AlleCompanion with Allegro’s production baselines on desktop and mobile web, grouped as Web, and in the mobile app. Existing systems combined item-to-item collaborative filtering, which suggests items that similar shoppers bought, with seller bestsellers, new arrivals or recurring purchases, depending on the placement.\n\nAllegro measured visit conversion, carousel conversion and GMV attributed to carousel interactions. The paper marks results significant at a threshold of p less than 0.005. It omits user counts, traffic allocation between variants, confidence intervals and absolute baseline values.\n\nOn the sponsored product-page module, complementary recommendations increased visit conversion by 0.53% on Web, a statistically significant result, and by 0.13% in the app, which fell short of significance. GMV barely moved. Allegro still reports that advertising revenue from the carousel rose by about 50% across both platforms because click-through improved. That revenue figure is company attribution, and the paper offers no uncertainty range for it.\n\nThe organic product-page module required a different policy. Pure complements had proved too restrictive because shoppers also wanted alternatives while browsing. Adding same-category products produced a reported GMV increase of 8.05% on Web and 9.35% in the app, both statistically significant, while visit conversion stayed flat.\n\nThe two pre-cart tests were the counterexample. The expert-filtered model and the pretraining and fine-tuning strategy both failed to improve GMV. Changes ranged from plus 0.17% to minus 0.28% across Web and app. The expert-filtered treatment also cut Web visit conversion by 0.28%, the only statistically significant movement in that test.\n\nShopping cart recommendations behaved differently one step later. Inside the cart, a mixed policy that included same-category alternatives increased attributed GMV by 21.25% on Web and 15.73% in the app. Carousel conversion rose 4.98% on Web and 1.21% in the app, with the Web result significant.\n\n## One click changed the commercial effect\n\nPre-cart and in-cart placements both sat close to checkout and both encouraged shoppers to add products from the same seller to reach a free-delivery threshold. Yet the same broad recommendation policy produced a flat GMV result in one location and a large attributed effect in the other.\n\nThe paper leaves the reason open, and any explanation about commitment, visibility or readiness to add another item would be inference. The comparison does show that offline relevance and the apparent similarity of two interfaces both failed to predict placement performance.\n\nContext therefore belongs in the model evaluation. The retrieval system, category policy, existing baseline and shopper’s current task formed the treatment together. A technically stronger model could win in one placement and stall in the next, so a single validation could never cover the whole journey.\n\nWithout the failed pre-cart result, the strongest figures could be read as proof that complementary product recommendations generally raise basket value. With it, the evidence supports a different lesson. The system created value where its policy matched the job of the placement.\n\n## Allegro deployed the three winners and shelved the rest\n\nAllegro deployed AlleCompanion to the sponsored product page, organic product page and in-cart modules. The stated rule required a statistically significant improvement in at least one primary metric on at least one platform, with the remaining metrics unharmed.\n\nThe model now serves those placements at a marketplace with more than 20 million monthly active users. The system covers 99.8% of active customer interactions, according to the authors. They still flag cold-start taxonomy gaps (new products that lack category history), limited direct personalization in candidate generation and feedback loops that make offline evaluation difficult.\n\nFor a European operator, the sequence offers more to copy than the size of the GMV lifts. Allegro inspected noisy labels before training, kept business policy outside the retrieval model, ran temporal offline tests, moved to placement-level online experiments, kept a failed result in the paper and deployed only where the full system cleared a commercial guardrail.\n\nThe Allegro recommendation system treated a good related product as something that changes with the moment. Allegro made that definition adjustable, then tested it where each recommendation appeared. The gap between flat GMV and a 21% attributed lift came down to one step in the customer journey, with the neural architecture held constant.\n\n[Pinterest’s image-search experiment](https://industrycontents.com/pinterest-image-search/) exposed the same reason to judge a retrieval system inside the live interface rather than rely on offline relevance alone.\n\n## Sources and limits\n\n- [Allegro’s AlleCompanion paper](https://arxiv.org/html/2609.05063v1) , released September 4, 2026\n- [Allegro’s official company overview](https://about.allegro.eu/)\n- [Lead author Aleksandra Osowska-Kurczab’s LinkedIn profile](https://pl.linkedin.com/in/aleksandra-osowska-kurczab)\n\nAllegro employees authored the paper, with one author at NVIDIA, the chipmaker, when it was published after completing the work at Allegro. Independent teams have yet to replicate the experiments. The paper reports relative changes and a significance threshold and omits sample counts, variant allocation, absolute baselines, confidence intervals and the exact test statistic. Its GMV measure is attributed to carousel interactions, so it describes carousel-driven sales and stays short of a platform-wide revenue lift.\n\nThe expert-filtered dataset removed 77% of the available training pairs. It improved attribute consistency and weakened standard relevance measures. Models trained only on synthetic expert-rule pairs also performed poorly. The strongest compromise used synthetic pretraining followed by fine-tuning on behavioural data.\n\nLabels that look cleaner to experts can remove the behavioural variation a model needs in production. Allegro resolved that tension by moving precision into the controllable category layer and leaving the retrieval model exposed to much of the real distribution.\n\n## Five tests put the same system in five contexts\n\nThe ecommerce recommendation experiment compared AlleCompanion with Allegro’s production baselines on desktop and mobile web, grouped as Web, and in the mobile app. Existing systems combined item-to-item collaborative filtering, which suggests items that similar shoppers bought, with seller bestsellers, new arrivals or recurring purchases, depending on the placement.\n\nAllegro measured visit conversion, carousel conversion and GMV attributed to carousel interactions. The paper marks results significant at a threshold of p less than 0.005. It omits user counts, traffic allocation between variants, confidence intervals and absolute baseline values.\n\nOn the sponsored product-page module, complementary recommendations increased visit conversion by 0.53% on Web, a statistically significant result, and by 0.13% in the app, which fell short of significance. GMV barely moved. Allegro still reports that advertising revenue from the carousel rose by about 50% across both platforms because click-through improved. That revenue figure is company attribution, and the paper offers no uncertainty range for it.\n\nThe organic product-page module required a different policy. Pure complements had proved too restrictive because shoppers also wanted alternatives while browsing. Adding same-category products produced a reported GMV increase of 8.05% on Web and 9.35% in the app, both statistically significant, while visit conversion stayed flat.\n\nThe two pre-cart tests were the counterexample. The expert-filtered model and the pretraining and fine-tuning strategy both failed to improve GMV. Changes ranged from plus 0.17% to minus 0.28% across Web and app. The expert-filtered treatment also cut Web visit conversion by 0.28%, the only statistically significant movement in that test.\n\nShopping cart recommendations behaved differently one step later. Inside the cart, a mixed policy that included same-category alternatives increased attributed GMV by 21.25% on Web and 15.73% in the app. Carousel conversion rose 4.98% on Web and 1.21% in the app, with the Web result significant.\n\n## One click changed the commercial effect\n\nPre-cart and in-cart placements both sat close to checkout and both encouraged shoppers to add products from the same seller to reach a free-delivery threshold. Yet the same broad recommendation policy produced a flat GMV result in one location and a large attributed effect in the other.\n\nThe paper leaves the reason open, and any explanation about commitment, visibility or readiness to add another item would be inference. The comparison does show that offline relevance and the apparent similarity of two interfaces both failed to predict placement performance.\n\nContext therefore belongs in the model evaluation. The retrieval system, category policy, existing baseline and shopper’s current task formed the treatment together. A technically stronger model could win in one placement and stall in the next, so a single validation could never cover the whole journey.\n\nWithout the failed pre-cart result, the strongest figures could be read as proof that complementary product recommendations generally raise basket value. With it, the evidence supports a different lesson. The system created value where its policy matched the job of the placement.\n\n## Allegro deployed the three winners and shelved the rest\n\nAllegro deployed AlleCompanion to the sponsored product page, organic product page and in-cart modules. The stated rule required a statistically significant improvement in at least one primary metric on at least one platform, with the remaining metrics unharmed.\n\nThe model now serves those placements at a marketplace with more than 20 million monthly active users. The system covers 99.8% of active customer interactions, according to the authors. They still flag cold-start taxonomy gaps (new products that lack category history), limited direct personalization in candidate generation and feedback loops that make offline evaluation difficult.\n\nFor a European operator, the sequence offers more to copy than the size of the GMV lifts. Allegro inspected noisy labels before training, kept business policy outside the retrieval model, ran temporal offline tests, moved to placement-level online experiments, kept a failed result in the paper and deployed only where the full system cleared a commercial guardrail.\n\nThe Allegro recommendation system treated a good related product as something that changes with the moment. Allegro made that definition adjustable, then tested it where each recommendation appeared. The gap between flat GMV and a 21% attributed lift came down to one step in the customer journey, with the neural architecture held constant.\n\n## Sources and limits\n\n- [Allegro’s AlleCompanion paper](https://arxiv.org/html/2609.05063v1) , released September 4, 2026\n- [Allegro’s official company overview](https://about.allegro.eu/)\n- [Lead author Aleksandra Osowska-Kurczab’s LinkedIn profile](https://pl.linkedin.com/in/aleksandra-osowska-kurczab)\n\nAllegro employees authored the paper, with one author at NVIDIA, the chipmaker, when it was published after completing the work at Allegro. Independent teams have yet to replicate the experiments. The paper reports relative changes and a significance threshold and omits sample counts, variant allocation, absolute baselines, confidence intervals and the exact test statistic. Its GMV measure is attributed to carousel interactions, so it describes carousel-driven sales and stays short of a platform-wide revenue lift.", "url": "https://wpnews.pro/news/allegros-ai-had-a-good-idea-it-just-picked-the-wrong-spot", "canonical_source": "https://industrycontents.com/allegro-recommendation-system/", "published_at": "2026-10-01 08:00:00+00:00", "updated_at": "2026-10-01 08:16:12.045038+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-products", "large-language-models"], "entities": ["Allegro", "AlleCompanion", "ComCat", "Complementary Categories Mapping"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/allegros-ai-had-a-good-idea-it-just-picked-the-wrong-spot", "markdown": "https://wpnews.pro/news/allegros-ai-had-a-good-idea-it-just-picked-the-wrong-spot.md", "text": "https://wpnews.pro/news/allegros-ai-had-a-good-idea-it-just-picked-the-wrong-spot.txt", "jsonld": "https://wpnews.pro/news/allegros-ai-had-a-good-idea-it-just-picked-the-wrong-spot.jsonld"}}