How Netflix Uses Multimodal AI to Personalize Every Movie Poster (And How to Build It) Netflix's recommendation research team published a paper on MAPS (Multimodal Asset Personalization at Scale), describing how foundation models and multimodal embeddings replaced five separate ID-based artwork ranking models with a single two-tower architecture that encodes image and user-context features. The system addresses cold-start for new titles and thumbnails and enables aesthetic preference transfer across titles, with personalized posters reported to lift click-through rates by up to 30%. Have you ever sat on the couch with a friend and realized that Netflix was showing you two completely different posters for the exact same movie? If your profile has a history of watching romantic comedies, Netflix’s homepage might show you a thumbnail of Good Will Hunting featuring Matt Damon and Minnie Driver sharing an intimate smile. If your friend spends their evenings watching high-octane action thrillers or stand-up specials, their homepage might show an alternate thumbnail highlighting Robin Williams looking thoughtful, framed like a character drama. This is not a visual gimmick. For Netflix, promotional artwork is the single largest influence on a subscriber’s decision to press play. A personalized poster can increase title click-through rates by up to 30%. Yet for years, the engineering architecture powering this visual personalization suffered from a massive, hidden flaw: the machine learning models were completely blind. They had no idea what was actually inside the image. They did not know whether a poster showed an explosion, a close-up of an actress, or a typography logo. They only knew random database IDs: user 8192 clicked artwork 304 . Last week, Netflix’s recommendation research team published a landmark paper on MAPS Multimodal Asset Personalization at Scale . They revealed how foundation models and multimodal embeddings allowed them to collapse five fragmented machine learning systems into a single, unified architecture that sees and hears every asset it recommends. Here is how Netflix rebuilt their visual engine, the mechanics of two-tower multimodal ranking, and the blueprint to implement it in your own stack. 1. The Blind Spot: The Cold-Start Disaster To understand why Netflix overhauled their architecture, you have to look at the limitations of traditional collaborative filtering . Historically, recommender systems treated personalization as a matrix factorization problem. You map users and items into low-dimensional latent spaces based purely on historical interaction logs. In production, this creates three painful engineering bottlenecks: 1. The Cold-Start Wall: When a new original film launches or a designer creates a fresh promotional thumbnail, the asset has zero click history. An ID-based model cannot recommend it because it has no data. The system has to fall back to random exploration, burning valuable homepage impressions. 2. Infrastructure Sprawl The 5-Canvas Problem : Netflix displays artwork across five distinct UI surfaces: the wide TV hero banner, the vertical mobile card, the row thumbnail, the search discovery grid, and the tablet carousel. Because interaction dynamics differ across devices, Netflix had to train, maintain, and monitor five separate ID-based models . 3. Contextual Blindness: An ID model cannot transfer learning across titles. If a user consistently responds to posters featuring close-up faces with dramatic warm lighting, an ID-based model cannot generalize that aesthetic preference to a newly added indie film. The solution was obvious in theory, but notoriously difficult at scale: teach the recommendation model to actually see the image. 2. The Architectural Shift: Two-Tower Multimodal Ranking Instead of relying on arbitrary integer IDs, modern recommender systems decouple candidate scoring into a Two-Tower Architecture : In this architecture, candidate scoring is split into two independent representation pipelines: The User Context Tower Left This tower encodes dynamic real-time user context: - Historical Taste Vectors: Genres watched over the past 30 days, preferred actors, and visual tropes the user historically engages with. - Canvas Metadata: The active UI surface mobile portrait vs. TV landscape banner . - Session Context: Time of day, device type, and regional language settings. The Multimodal Asset Tower Right Instead of an asset ID, this tower ingests rich, high-dimensional foundation model embeddings: - Visual Embeddings CLIP / SeqCLIP : Pretrained vision-language encoders convert raw image pixels and video frames into a normalized 512-dimensional semantic vector, capturing lighting, facial expressions, color palette, and composition. - Audio Embeddings wav2vec 2.0 : For personalized video preview clips, an audio model encodes background music, sound effects, and emotional tone. - Text & Subtitle Embeddings: Time-aligned dialogue transcripts captured via lightweight sentence transformers. Real-Time Dot-Product Ranking At inference time, scoring an asset does not require running heavy vision models on the fly. All asset embeddings are precomputed offline and indexed in a vector store. The user tower outputs a context vector, and a fast dot-product operation computes the cosine similarity across candidate posters in sub-50 milliseconds. Because the model scores images based on their visual semantic features rather than their database ID, a newly uploaded thumbnail can be personalized on day one: zero-shot cold start is solved. Even better: because the model understands the visual properties of the asset, a single two-tower model now serves all five UI canvases , permanently retiring five legacy models. 3. Query-Aware Artwork: Personalization in Search The most elegant byproduct of adopting multimodal embeddings is what happens when a user navigates to the search bar. In traditional search engines, if you type “comedy” , the search engine retrieves comedy titles and displays the standard, generic title poster. Because Netflix’s multimodal system uses a joint text-image embedding space CLIP , artwork selection becomes dynamically query-aware: Here is how it works under the hood: 1. The user types a natural-language search query: “dark psychological thriller” . 2. The query is passed through CLIP’s text encoder, producing a 512-dimensional query vector. 3. For every retrieved title matching the query, the engine evaluates all available candidate thumbnail frames in that movie’s asset pool. 4. The system calculates the dot-product similarity between the text query vector and the image frame vectors . If a movie contains both comedic dialogue and dramatic suspense scenes, searching for “thriller” dynamically elevates the darkest, highest-contrast frame, while searching for the lead actor elevates a clean, recognizable close-up. The user sees visual evidence that the movie matches their specific intent before they even read the synopsis. 4. The 3 Production Scars: Lessons from the Real World Deploying foundation model embeddings into production recommender systems introduces three difficult systems bottlenecks: Scar 1: The Real-Time Serving Latency Trap Raw CLIP embeddings are typically 512 to 768 dimensions. If a user homepage has 40 title rows, and each title has 10 candidate thumbnails, computing real-time vector similarities across 400 candidates on every page load introduces severe p99 latency spikes. - The Guardrail: Netflix applies dimensionality reduction PCA / linear projection to compress embeddings down to lower-dimensional spaces e.g. 64 or 128 dimensions and precomputes candidate asset pools offline, caching them on fast edge key-value stores. Scar 2: The A/B Testing Gating Bottleneck Evaluating whether a new checkpoint of an embedding model improves recommendations usually requires spinning up an expensive multi-week online A/B test with live traffic. - The Guardrail: Netflix developed an offline proxy task : predicting the historical popularity-based winner using embeddings alone. If an experimental model cannot beat the baseline embedding checkpoint on this offline ranking proxy, it is discarded before it ever touches live production traffic. Scar 3: Visual Dominance Over Audio in Video Personalization When Netflix trained MediaFM , their tri-modal foundation model for video preview clips, they found that visual features easily overpowered audio and subtitle signals during naive vector concatenation. The model was effectively ignoring the background score and dialogue. - The Guardrail: Multi-stage attention fusion. Instead of concatenating vectors, the architecture uses a cross-attention layer that forces the visual stream to condition on the audio emotional tone and spoken dialogue cues. 5. How to Prototype This Weekend Python Blueprint You can build a functional Two-Tower Multimodal Ranker in under 35 lines of Python using open-source tools: python Minimal Two-Tower Multimodal Ranker Blueprint import numpy as np class MultimodalArtworkRanker: def init self, embedding dim=128 : self.embedding dim = embedding dim Simulated learned projection matrix from user features to visual space self.user projection weights = np.random.randn 64, embedding dim def encode user context self, user feature vector: np.ndarray - np.ndarray: User Tower: Projects user viewing history into shared semantic space projected = np.dot user feature vector, self.user projection weights return projected / np.linalg.norm projected def rank candidate artwork self, user vector: np.ndarray, candidate clip embeddings: dict - str: Asset Tower: Dot-product similarity across precomputed CLIP visual embeddings best candidate = None highest similarity = -float "inf" for asset id, embedding in candidate clip embeddings.items : Normalized cosine similarity norm embedding = embedding / np.linalg.norm embedding score = np.dot user vector, norm embedding if score highest similarity: highest similarity = score best candidate = asset id return best candidate Example usage: ranker = MultimodalArtworkRanker embedding dim=128 User with a strong affinity for action simulated 64-dim vector mock user history = np.random.randn 64 user context = ranker.encode user context mock user history Precomputed CLIP embeddings for 3 candidate movie posters candidate posters = { "poster dramatic close up": np.random.randn 128 , "poster action car chase": np.random.randn 128 , "poster romantic embrace": np.random.randn 128 } selected poster = ranker.rank candidate artwork user context, candidate posters print f"Selected Thumbnail: {selected poster}" 6. The Strategic Bottom Line For technical leaders, Netflix’s MAPS project reveals where the next generation of recommender systems is heading: Content-blind recommenders are obsolete. Relying solely on historical interaction IDs leaves your platform helpless whenever new products, fresh creative assets, or unrated inventory enter your catalog. By anchoring your recommendation engines to multimodal foundation embeddings: 1. You solve the cold-start problem on day one. 2. You collapse infrastructure sprawl from multiple device-specific models into a single unified engine. 3. You unlock query-aware visual adaptation across search and discovery surfaces. Personalization isn’t just about predicting which title a user wants to watch. It is about presenting that title through the visual lens that resonates with them the most. Further Reading & Resources - Netflix Research: Multimedia Asset Personalization via Multimodal Embeddings https://arxiv.org/abs/2608.18322 : The original technical paper by Aditya Deshpande et al. detailing the MAPS architecture and MediaFM foundation model. - OpenAI: Learning Transferable Visual Models From Natural Language Supervision https://arxiv.org/abs/2103.00020 : The foundational CLIP paper explaining joint vision-language representation spaces. - Meta: wav2vec 2.0 Framework https://arxiv.org/abs/2006.11477 : Self-supervised learning of speech and audio representations utilized in multimodal video models. - Netflix TechBlog: Artwork Personalization at Scale https://netflixtechblog.com/artwork-personalization-c589f074ad76 : Historical background on how contextual bandits paved the way for modern embedding-based asset ranking. If you enjoyed this breakdown, subscribe to MLnotes https://mlnotes.substack.com/ for weekly, bite-sized systems engineering and AI architecture deep-dives. If your product team is building recommendation engines, share this article with them.