cd /news/artificial-intelligence/generative-ai-for-recommendations-wh… · home › topics › artificial-intelligence › article
[ARTICLE · art-144828] src=pub.towardsai.net ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Generative AI for Recommendations: what YouTube, Netflix and Meta are Moving to, Built From Scratch

A tutorial builds a generative AI recommender from scratch on the Amazon Reviews 2023 dataset from McAuley Lab (UCSD), using the Video Games category in its 5-core version where every user and item has at least 5 interactions, and measures what each component adds over classic collaborative filtering. The post notes that up to 28% of test users were never seen during training, and that 94% of users in the raw data wrote exactly one review, motivating generative recommendation that treats a user's history as a sentence and generates the next item from item content such as title and description. YouTube, Netflix and Meta have all published work moving their recommendations in this direction.

by read19 min views1 publishedOct 4, 2026

The problem they’re solving is pretty simple to state: there are too many items (products, movies, videos, songs) for anyone to browse through all of them, so something has to decide what’s worth showing each person. Get it right, and people find things they actually want. Get it wrong, and it’s just noise.

The core idea behind this kind of recommendation is actually pretty intuitive: if two people have liked similar things in the past, they’ll probably like similar things in the future too.

The algorithm’s job is to figure out, mathematically, who is similar to whom and what is similar to what, without anyone telling it what “similar” means. That is the classic, collaborative way of recommending: compare people, and compare items, through who interacted with what. A generative recommender does something different. It treats your history as a sentence and writes the next item, the way a language model writes the next word and because each item’s “word” is built from its content, it can write items nobody has bought yet.

This works well until something new shows up. A new user has no history, and a new product has no buyers, so the model has nothing to compare. In the dataset I used, up to 28% of test users had never been seen during training.

That’s where a generative recommender comes in. Instead of scoring items by similarity, it treats your history like a sentence and generates the next item, the way a language model generates the next word. And because each item’s “word” is built from its content (its title and description), a brand-new product can be recommended before anyone has bought it.

This isn’t a research curiosity: YouTube, Netflix and Meta have all published how they’re moving their recommendations in this direction.

In this post I’ll build a generative AI recommender from scratch, on real Amazon data, and measure what each piece actually adds over the classic approach. When you finish this blog, you can understand how it works and build yourself one too.

I used the Amazon Reviews 2023 dataset from McAuley Lab (UCSD), the same family of data the YouTube and Google papers benchmark on. Specifically the Video Games category, in its “5-core” version: every user and every item has at least 5 interactions.

Why 5-core? Because a recommender learns from history, and the raw reviews barely have any: in the raw data, 94% of users wrote exactly one review. You can’t predict “what’s next” for someone with no “before”.

Let’s start with the simplest baseline and add complexity from there. First, the popularity model.

The idea of this algorithm is to count how many times each item was interacted with, across all users. Sort from most to least popular. Cross off anything this user already has. Whatever’s left at the top is the recommendation.

Notice that everyone starts from the same ranked list. The only personalization is the crossing off: in the example, D is the most popular item by far, but this user already has it, so it’s removed and B, E and H become the top 3.

In math, the score of an item is just its count (over all users):

and the recommendation for user u is the top-K by that score, skipping the items in their history H_u:

Popularity gives everyone the same list. To personalize, we need a way to represent what each user tends to like, without anyone defining what “tends to like” means.

Start with a table. Put users in rows and items in columns. Each cell says whether that user interacted with that item: 1 if yes, and a blank if we don’t know. Call this matrix R. Recommending is filling in the blanks !

Now we will say this R matrix is built coming from an interaction of two matrices X and Y. Each user gets a row of X, call it x_u, and each item gets a row (*** r_ui***) of Y, call it y_i. Both have the same length k (2 in the figure).

Where do they come from? X and Y start as random numbers. Then the model checks how well X·** Yᵀ** reproduces the 1s actually observed in R, adjusts the numbers to do a bit better, and repeats.

The prediction for any cell is those two rows multiplied position by position and added up.

We can see in the matrix R, the user 1 and item 2 never interacted, the decomposition X·** Yᵀ** should produce a score for it, while user 1 and item 1 did interact, naturally will produce a similar number for it too.

How do we pick the numbers? We want predictions that match the cells we know. So for every known cell, take the gap between the real value and the prediction, square it (so negative and positive gaps don’t cancel), and add all the gaps up. That total is the minimization problem we want to solve. Smaller loss, better numbers for recommending:

If you follow up to here, you understand the algorithm. What comes next are two refinements that make it work better in practice (feel free to skip them). The minimization problem has far more freedom than data. It can drive the loss to zero by making the numbers huge and oddly shaped, fitting the cells perfectly and predicting nonsense everywhere else. That’s overfitting. The fix is to add a penalty on the size of the numbers (Regularization), so the model only grows them when the data really justifies it

λ controls the trade-off: bigger λ, smaller numbers, smoother predictions.

We never see a “no”. A blank means “not observed”, not “disliked”. The fix from Hu, Koren and Volinsky (2008) is to trust the 1s a lot and the blanks only a little, with a confidence weight on every cell:

Putting the weight into the loss gives the final version, the one the implicit library implements:

Solving for X and Y at the same time is hard: the loss has products of unknowns in it. The trick is to freeze one side. If Y is fixed, each user’s row x_u appears only inside squared errors, so it’s a plain least-squares problem with an exact solution. Freeze X instead, and the same is true for each item. Now if we fixed from the minimization (consider as constant) the other expression, the minimization (dL/dx=0 or dL/y=0) (proof avoided for complexity, non trivial) solution becomes:

The data is split by time, not at random: everything before November 2019 is training, November 2019 to October 2021 is validation, and October 2021 to September 2023 is test. A random split would let a model train on a 2022 review and be tested on one from 2005 it would have seen the future

This is what Recall@K answer: did we find it? Of the items the user actually interacted with, what fraction appear anywhere in the top-K list? Order inside the list doesn’t matter.

If the user bought two items and one of them is in our top 10, Recall@10 = 0.5. This is what NDCG@K: did we find it early? Recall treats a hit at position 1 and a hit at position 10 the same. But a user scrolls from the top, so position 1 is worth more. NDCG gives each hit a weight that shrinks with its position, then divides by the best score possible so the result lands between 0 and 1:

Both are averaged over all evaluation users, at at K = 5, 10 and 20 depending on the table.

One thing to know before the tables: the absolute numbers will look small a Recall@10 of 0.02 means 2%. Part of that is the dataset (25,612 items, a median of 6 interactions per user). But most of it is the question this split asks. Splitting by time means a user’s training history ends in 2019 and their test items are what they reviewed in 2021–2023 a median gap of almost four years, and 72% of the test items didn’t exist in the training data at all.

That is the realistic production question (“what will this person want next year?”), and it is a brutal one. The papers ask an easier and more standard one, and later in the post I measure that way too so the two can be told apart. Until then, what matters is the comparison between models, not the raw value.

You can play with this section in the notebook below:

Step 1: turn the title into numbers. A pretrained sentence model (sentence-t5-base, the same family TIGER uses) reads the product title and outputs a vector x of 768 numbers. Titles that mean similar things land close together in that space.

Step 2: compress it. A small encoder E squeezes those 768 numbers down to 32:

Step 3: snap it to a codebook. Instead of keeping z as 32 free numbers, we snap it to the nearest of 256 learned reference vectors, the “codebook” C₁. The item’s first digit is just the index of that nearest vector:

One snap is coarse. So we take what’s left over, the residual, and snap that to a second codebook, then repeat once more:

The Semantic ID is the triple (c₁, c₂, c₃). Three digits from 0 to 255: coarse, finer, finest. The item’s compressed vector is rebuilt by adding the three snapped vectors back together:

Three codebooks of 256 give 256³ ≈ 16.7 million possible IDs from only 768 learned vectors. And because items snap to the same first digit when their titles are close, the first digit acts like a category the model discovered on its own.

A tiny example, in 2 dimensions. Say z = (0.72, 0.31) and the first codebook has three entries: (0.2, 0.8), (0.8, 0.2), (0.5, 0.5). The nearest is (0.8, 0.2), so c₁ = 2. The residual is (0.72 − 0.8, 0.31 − 0.2) = (−0.08, 0.11). The second codebook has (−0.1, 0.1), (0.1, −0.1), (0, 0); the nearest is (−0.1, 0.1), so c₂ = 1. The ID is (2, 1), and the reconstruction is (0.8, 0.2) + (−0.1, 0.1) = (0.7, 0.3), off from z by only (0.02, 0.01).

Step 4: train it. A decoder tries to rebuild the original 768 numbers from z. The whole thing is trained to make that rebuild accurate, while also pulling the codebook vectors toward the data and the encoder toward the 3 codebooks:

sg[·] means “stop gradient”: the middle term moves only the codebooks, the last one moves only the encoder. Without the last term the encoder drifts and never commits to a code.This is what we call (Residual Quantization) RQ-VAE.

What came out. After 400 epochs the loss went from 3.4 to 1.03. Of 25,612 items, 21,713 got a unique 3-digit code; the remaining 25% shared a code with at least one other item, so a fourth digit is appended just to tell them apart, exactly as TIGER does. Levels 2 and 3 used 96% of their 256 codes; level 1 used 29%, which makes sense: there aren’t 256 kinds of video-game product. A concrete sign it worked: five items that landed on the same first digit were five gaming keyboards and mice, a grouping the model was never told to make.

The embedding step alone already knows what things are. Nearest neighbours of one product title, by cosine similarity in the Sentence-T5 space. Nothing about “headset” was ever labelled; that structure is what the RQ-VAE compresses into three digits. Check next google colab to interact with this process !

The classic way to recommend is search: embed the user, embed every item, find the closest ones. The generative way, from TIGER and Meta’s HSTU, is to write the next item’s ID directly, digit by digit, the way a language model writes the next word. Predicting “what’s the next item” becomes the exact same problem as predicting “what’s the next word”.

Step 1: the history becomes a sentence. Take a user’s purchases in time order and replace each item by its four digits (three Semantic ID levels plus the tie-break digit). Twenty items become a stream of 80 tokens, plus one token at the front that identifies the user:

Nothing in that stream says “product” or “user”. It’s just tokens, and tokens are what language models eat.

Step 2: train a small language model on it. A decoder-only Transformer — the architecture behind GPT, hand-written here rather than imported: 4 layers, 4 attention heads, 128-dim, dropout 0.1, about 1.1M parameters — reads the stream and, at every position, predicts the next token from everything before it. Training maximises the probability of the token that actually came next, which is the same as minimising cross-entropy:

That’s the whole training objective. There’s no “similarity”, no user–item matrix: the model just learns which digits tend to follow which. Training took 45 passes over the data on a Colab GPU, bringing the loss from 4.2 (random digits) to 2.2 per digit and it was still going down slowly when I stopped, which will matter later.

Step 3: generate. At inference, feed the user’s history and let the model write four more tokens. The probability of a full item is the product of its four digit probabilities, each conditioned on the digits written so far:

The first digit picks the broad family, the second narrows it, the third pins the item down, the fourth breaks ties. Similar items share prefixes, so the model can put probability on a whole family at once including items it has never seen anyone buy.

Step 4: only write real items. Not every 4-digit combination is a product. So generation is constrained: at each step the model may only choose digits that lead to an existing item, checked against a lookup tree built from the catalog. And instead of greedily taking the single best digit each time, beam search keeps the 30 best partial IDs alive and extends each one. The score of a finished candidate is its log-probability; drop anything the user already has, and the top-K by score are the recommendations:

Every recommendation this produces is guaranteed to be a real item, nothing needs filtering afterward.

One concrete case. One user’s last purchases were all Nintendo Switch accessories: a controller, a case, a stand, a charger. The model’s top 10 were all Switch products, and #1 was exactly what that user bought next. It was never told the platform; it picked that up from the pattern of digits in the history.

Retrieval has one job: be cheap and wide. ALS scores all 25,612 items with a single matrix multiply; the Transformer writes candidates digit by digit. Neither can afford to look at price, or at how recently an item has been selling, or at real content similarity, for every item every time. So production systems add a second stage: take the ~100 candidates retrieval already produced, and spend a bit more compute per candidate to put them in a better order. Same shortlist, better order.

The features. For each (user, candidate) pair we compute seven numbers: retrieval’s own confidence score, how popular the item is overall, its price (with a flag for the 25% of items where the price was missing and had to be filled in), how recently it’s been trending, how long the user’s history is, and, the one that ties back to the Semantic ID work, the cosine similarity between the item’s title embedding and the average embedding of the user’s past purchases. The same embedding space that got quantized into Semantic IDs, used continuously here.

The model. A gradient-boosted tree ensemble (LightGBM) that maps those seven numbers to a score:

What makes it a ranker rather than a classifier is the objective. LambdaMART doesn’t ask “is this item a hit, yes or no?”; it looks at pairs. For every pair where item i was the true next purchase and item j was not, it pushes s_ui above *** s_uj***, and it pushes harder when swapping the two would change NDCG more, i.e. when the mistake is near the top of the list:

So the model is trained directly on the metric we report. The final list is just the candidates sorted by score:

Training data. 18,000 users, ALS’s top-100 candidates for each, labelled by what they actually bought next. Only 2,347 of those users had their true next item anywhere in ALS’s top 100 that is retrieval’s recall@100, and it is the ceiling a reranker works under. Those 2,347 users, 235K rows, 2.9K positives, are the training set. Training took under a minute.

The bug I shipped first. My first version “fixed” the missing positives: when the true item wasn’t in ALS’s top 100, I injected it into the list with a placeholder retrieval score (the lowest one in the list). Since ALS misses the true item about 90% of the time, 93% of the ranker’s positives were injected and it learned exactly what I had taught it: the lowest-scored, never-seen candidate is the one to promote. Probing the trained model with every other feature held fixed, a retrieval score of 0.1 got +0.4 and a score of 0.5 got −4.4; an item with zero training interactions got +6.5. On real candidates that is the opposite of useful, and the reranker duly broke even with plain ALS. A reranker can only reorder what retrieval gives it, so it has to be trained on what retrieval gives it, with the real feature values. That is the whole fix.

What the trees learned to look at, after the fix: retrieval’s own score first, then content similarity , the Semantic-ID embedding space, used continuously, then how recently the item has been trending, then popularity. Price and the user’s history length mattered least.

Every tier: popularity, ALS, the Semantic-ID Transformer, ALS plus the reranker scored on the same 2,000 users per split under the time-based split, with the same rule for the users who have no training history at all (everyone falls back to the popularity list for them; that is 29–42% of the sample, and excluding them would be hiding the hardest part of the problem).

Two things stand out, and one of them bothered me for a while.

The reranker is the clear winner: 35–50% over the best retrieval-only tier on every metric, both splits. That is the expected shape of a two-stage system: retrieval is cheap and wide, ranking spends the compute where it counts.

The other is that the three retrieval tiers are all at the same floor, and the generative model — the whole point of the post — does no better than counting what’s popular. Is it broken? I spent a while on that question, and the answer turned out to be no. It’s the split.

Look again at what the time split asks: from a median of four reviews written before November 2019, predict which games this user reviews in 2021–2023. The median gap between a user’s last training interaction and their first test interaction is 1,397 days, and 72% of the test items never appear in training. No collaborative model can retrieve an item it has never seen, and no model of any kind is good at four-year-ahead prediction from four data points. Everything collapses to the same floor, and the floor is “recommend what’s popular”.

TIGER and SASRec and essentially every sequential-recommendation paper measure something different: leave-one-out next-item prediction. For each user, the last item is the test target, the second-to-last is the validation target, and the context is everything before. The context ends right before the target and the target almost always exists in training (0.1% cold instead of 72%). Same data, different question, and the numbers are an order of magnitude apart.

So I added that protocol as a second track, and with it the baseline the papers actually compare against: SASRec (Kang & McAuley 2018), which is the same causal Transformer over raw item IDs instead of Semantic IDs, scoring the next item by a dot product with the item table. It’s the cleanest possible control for the question the post is really about. Does writing items out of content-derived digits beat looking them up by ID?

All 94,762 users, both splits, each user’s own history filtered from every model’s list:

First: the generative model works. Under the paper’s protocol it reaches Recall@10 of 0.069–0.077, nearly three times the popularity floor and in the range TIGER reports on its own datasets (0.065 on Amazon Beauty). The earlier table was measuring a task nothing here can do well, not a broken model.

Second: it does not beat SASRec on this dataset. SASRec is ahead by about 15% on every metric, where TIGER’s paper reports the opposite. My list of suspects, cheapest to test first: the Transformer was still improving when I stopped (TIGER trains for ~200K steps; I did a few thousand); the RQ-VAE only uses 29% of its first-level codebook, so the Semantic IDs are a weaker partition than they should be; the model is decoder-only where TIGER uses an encoder–decoder; and Video Games has a heavy head of blockbuster titles and consoles that an ID-based model can simply memorise, while the content-based model has to reach them through shared digits. Each of those is a concrete experiment, and that is the point of having the table.

The evaluation protocol is part of the result. Same data, same models, and the ordering of the tiers flipped between “predict next year” and “predict the next item”. Neither is wrong; they answer different questions, and a post that only shows one of them is hiding something.

A reranker trained on the wrong candidates learns the wrong thing, confidently. The injection bug produced a model with sensible feature importances and a plausible score, and it was learning the exact inverse of what it should. Probe your model’s response to each feature; it takes five minutes.

Generative retrieval from content-derived IDs is real: on the standard protocol it lands where the papers say it should, and it can name items nobody has bought. On this dataset it still loses to a well-trained ID-based Transformer, and the gap has a shortlist of known causes rather than a mystery. That’s a better place to end than a win I couldn’t explain.

Code, data pipeline and every table in this post: github.com/juanmigutierrez/generative-recommendation-engine. Each section has an “Open in Colab” notebook that runs a small version of the same step in under ten minutes.

Architecture

Building blocks

Baselines and ranking

Data

Software: implicit (ALS), sentence-transformers (Sentence-T5 inference), JAX (RQ-VAE, Transformer, beam search), LightGBM.

Code and notebooks: github.com/juanmigutierrez/generative-recommendation-engine. Generative AI for Recommendations: what YouTube, Netflix and Meta are Moving to, Built From Scratch was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @youtube 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/generative-ai-for-re…] indexed:0 read:19min 2026-10-04 · —