GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model Meta's Generative Ads Recommendation Model (GEM), the foundation model behind ads recommendations on Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs, achieving a doubling of end-to-end training efficiency to 20–25% Model FLOPs Utilization (MFU) while scaling training FLOPs 4x, as detailed in a post on Engineering at Meta. Meta’s Generative Ads Recommendation Model GEM , the foundation model behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs. This post goes into the details on how we achieved: doubling end-to-end E2E training efficiency to 20–25% Model FLOPs Utilization MFU while scaling training FLOPs 4x in ... The post GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model https://engineering.fb.com/2026/08/03/ml-applications/training-gem-at-llm-scale-meta-ads-recommendation-foundation-model/ appeared first on Engineering at Meta https://engineering.fb.com .