I would not start with a general-purpose LLM for this. Video recommendation is mainly a retrieval-and-ranking problem, and the best model depends on what interaction data you have.
A practical open-source starting point is TensorFlow Recommenders (TFRS) with a two-tower retrieval model:
The official movie-retrieval tutorial is close to this use case and includes training, evaluation, and export: https://www.tensorflow.org/recommenders/examples/basic_retrieval
If you have ordered watch histories, compare sequential models such as **BERT4Rec** or SASRec through RecBole. BERT4Rec is specifically designed to model sequences of user interactions: [BERT4Rec — RecBole 1.2.1 documentation](https://recbole.io/docs/user_guide/model/sequential/bert4rec.html)
For a cold-start baseline, you can embed titles, descriptions, genres, and tags with a text encoder such as `sentence-transformers/all-MiniLM-L6-v2`
and recommend semantically similar videos. Its model card lists semantic search as an intended use: sentence-transformers/all-MiniLM-L6-v2 · Hugging Face . However, that is an item-content encoder, not a complete personalized recommender.
My suggested progression would be:
So, if I had to choose one free starting stack: TFRS two-tower retrieval plus a ranker, with content embeddings for cold start. I would choose BERT4Rec only after confirming that sequential watch history improves the offline metrics.