Build an Art Recommender from scratch with CLIP and ChromaDB A three-part tutorial series shows developers how to build an art and museum recommender from scratch using CLIP embeddings and ChromaDB, with part one covering data retrieval from the Rijksmuseum API. The author, writing from personal experience as a junior data science practitioner, notes that museums including the Art Institute of Chicago and the Metropolitan Museum of Art have restricted API requests because agents are flooding public endpoints, leaving the Rijksmuseum as a well-documented alternative. The series promises a working laptop prototype, with part two covering vector database construction and part three the adaptive recommender system. A couple of years ago, as I was taking my first steps in the data science world, I had an idea and decided to see if I could build a full-fledged art and museum recommender with no prior knowledge. Yes, none . It was a challenge. I had never set up a database, didn’t know how to collect, what do with the data if I had it, or what were the strategies to provide recommendations. I was starting from scratch. And that is why this project taught me so much. Looking back, it was one of the best learning experiences I’ve had in my brief career so far. I believe this kind of project could be a great opportunity for anyone wanting to move from junior ML practitioner to a more experienced engineer, with end-to-end competencies and knowledge of recommender systems. And as a bonus, admire beautiful art along the way. That is why I’ve decided to write this three-part account of the project. If you follow it closely, I can assure you that you will end up with a better sense of how recommender systems work, and you’ll have a working prototype on your laptop. Since I want to take you through the entire project, this is going to be part one of three. Part two will explore how to turn your images into a machine-readable format and build a vector database with ChromaDB, while part three will put it all together to create an adaptive recommender system. If this caught your attention, read on The first thing I learnt when I decided to start with this project was that art data is ideal: easily available, well-curated, free, and mostly copyright-free What more could a data scientist want? It turns out that several museums provide data from their collections as a service to the general public. So if you want to build something on top of it, you just go and connect to their API. And if you have lots of high-quality data, sky is the limit. Since then things have changed unfortunately. Agents are running wild on the web and any public API is flooded with requests. This is why museums like ARTIC https://www.artic.edu or the MET https://www.metmuseum.org now restrict requests via their API. Luckily for us, there are still some museums providing this data via API, and one of them is the Rijksmusem https://www.rijksmuseum.nl/en in Amsterdam, which just so happens to be one with a beautiful collection. So now that we know the data exists, and that the Rijksmuseum API https://data.rijksmuseum.nl is well-documented, let’s get some data. After playing a little bit with the Rijksmuseum API, as well as those from other museums, it made sense to build a reference class that contained most of the important logic to download and do a first selection of the artworks to be used for the project. Some of the requirements were: The GenericMuseumApiInteractor class https://github.com/SimonCalo/art-recommender-from-scratch/blob/main/data retrieval/src/generic museum api interactor.py creates the blueprint for these functionalities. But let’s take a closer look at some of the key points. php def get response dict self, url: str, n trial: int = 0 - dict: json data = {} try: include a timeout to avoid getting stuck on one request response = requests.get url, timeout=10 Check if the request was successful if response.status code == 200: json data: dict = response.json else: print f"Error: {response.status code} - {response.text}" if the request timed out, try again for a max. of 5 times except requests.Timeout: if n trial < 5: print f"Trial {n trial} failed. Sending request again" self.get response dict url=url, n trial=n trial+1 else: print "Unable to reach link" return json data This first function is the one used to get the data from the museum API. Some relevant aspects are: Another valuable operation is removing duplicates: php def download image self, url: str, image name: str, n trial: int = 0 - bool: try: response = requests.get url, timeout=10 Check if the request was successful if response.status code == 200: Get the content of the response image data = response.content if ".jpg" in str image name : image name = image name.rstrip ".jpg" Open the image to check for duplicates img data = Image.open io.BytesIO image data image hash = imagehash.phash img data if image hash not in self.hashes: self.hashes.append image hash Specify the local file path where you want to save the image local file path: str = f"{self.museum name}/images/{image name}.jpg" Save the image data to the local file with open local file path, "wb" as image file: image file.write image data ... Using imagehash we check whether the image is identical to one we’ve already downloaded. It’s better to do this as early as possible to avoid polluting your data. No one likes to be recommended the same image five times. The last segments of this class worth checking are: php def is exception self, artwork dict: dict - bool: and php def passes requirements self, artwork dict: dict, is nested: bool = False - bool: simply because it’s better to do a bit of curation upfront to remove data that would reduce the quality of your final product. These two functions help with selecting only public domain images, as well as removing, for instance, fragments of artefacts. All of these aspects can be controlled via museum-level configurations. Here https://github.com/SimonCalo/art-recommender-from-scratch/blob/main/data retrieval/config/apis params config.py is an example of how these can be defined: museums details = { "met": { "base url": "https://collectionapi.metmuseum.org/public/collection/v1", "key": None, "exceptions": { "isPublicDomain": False, "objectName": "fragment" , "title": "fragment" , }, "requirements": {}, "results per page": None, }, "rijks": { "base url": "https://data.rijksmuseum.nl/search/collection?imageAvailable=true", "key": None, "exceptions": {}, "requirements": { "subject to 0 .classified as 0 . label": "public domain", }, }} In general, I would recommend using YAML files to control configurations; they are more light-weight and standardised. Here, I have used a .py file for simplicity. The idea is to create a structure that covers the most important configurations needed to interact with the API of different museums. Having created this generic blueprint class for interacting with museum APIs, we can move on to create a specific class https://github.com/SimonCalo/art-recommender-from-scratch/blob/main/data retrieval/src/rijks api interactor.py to interact with the Rijksmuseum API: python from data retrieval.src.generic museum api interactor import GenericMuseumApiInteractorclass RijksApiInteractor GenericMuseumApiInteractor : def init self - None: super . init museum name="rijks" def run downloading pipeline self, n images to download: int - None: super .run downloading pipeline page url = self.base url Loop over the pages while n images to download 0: structure the query parameters in the correct format extract the data for all the images on this page page response: dict = self.get response dict url=page url if page response == {}: continue go through every result in the page for item in page response "orderedItems" : artwork id: str = item "id" .split "/" -1 artwork dict: dict = self.get response dict url=item "id" if artwork dict == {}: print f"Artwork {artwork id} not found" continue visual item dict = self.get response dict url=artwork dict "shows" 0 "id" if visual item dict == {}: continue if not self.passes requirements visual item dict, is nested=True : continue digitally shown id = visual item dict "digitally shown by" 0 "id" image dict = self.get response dict url=digitally shown id image url = image dict "access point" 0 "id" if image url is None or image url.strip == "": continue image downloaded = self.download image url=image url, image name=artwork id if image downloaded: self.download json response dict=artwork dict, image name=artwork id n images to download -= 1 if n images to download == 0: break go to the next page page url = page response "next" "id" The main reason to create specific classes is because every museum has a different design in how their API is structured. In the Rijks one, for instance, results are displayed per page, and only by moving to the next page you get to access more artworks. Elsewhere, you might encounter all results to be displayed as a list of ID’s you need to access one by one. At this point, we are essentially done Just go to the root directory of your project or in a notebook , run the following: python from data retrieval.src.rijks api interactor import RijksApiInteractorinteractor = RijksApiInteractor interactor.run downloading pipeline n images to download=100 and watch the magic happen in front of your eyes. You should see images being downloaded with their corresponding metadata as json files automatically Another good option to retrieve images and related data is the Wikimedia Commons API https://commons.wikimedia.org/wiki/Commons:API . I haven’t tested it myself, but I’m sure it would add a lot of value after seeing other projects that have used it. Note: This part is not crucial to be able to proceed further. So if you’d rather dive and learn how to build the core of your recommender engine, feel free to skip directly to the end. There is one more aspect that is good to address for this kind of project; and that is data consolidation . If you plan to retrieve data from multiple sources especially public APIs that data will be structured in different ways. Some APIs might include dates as a list, others as two separate fields. If you want to work across these sources, you need to standardise the data you retrieve. This is made even harder by the fact that these API structures change constantly. So be prepared to refactor your code regularly When I was working on this project, I found that this kind of structure worked well for me: final metadata template = { "original id": "", "title original language": "", "title english": "", "artists": , "main artist": "", "artists details": , "museum api from which it was retrieved": "", "location": "", "year start number": -99999, "year start": "", "year end number": -99999, "year end": "", "dating of first display": "", "production places": , "main production place": "", "dimensions": , "artwork style": , "artwork subjects": , "artwork materials": , "artwork themes": , "artwork techniques": , "artwork types": , "artwork categories": , "subject description original language": "", "short description original language": "", "long description original language": "", "other descriptions original language": , "subject description english": "", "short description english": "", "long description english": "", "other descriptions english": , "colours": , "normalised colours": , "location within museum": "", "department within museum": "", "date processed": "", "documents and publications": } In general, select a few fields that you are interested in, such as title, description, author, etc. You can always add more later. Once you have created your ideal structure, you can create standardised json files for your metadata. I recommend using an LLM to help you with the logic, because looking across multiple json files to create a standard structure, finding the correct mapping of every field etc., is tricky. LLMs are much better than us at these kinds of tasks. Just feed it a couple of the jsons you have downloaded and the final structure you want, and let the tool do the work for you. Congratulations You have now laid the foundations on top of which you’re going to build your recommender. I am aware this first part might not have been the most exciting, but your recommender is only as good as the artworks it is built upon. And now, you have high-quality, open-access and curated data that you can use to construct your database. In the next part, we’re going to turn our data into vectors. Let’s get building Build an Art Recommender from scratch with CLIP and ChromaDB https://pub.towardsai.net/build-an-art-recommender-from-scratch-with-clip-and-chromadb-d6c948a7b979 was originally published in Towards AI https://pub.towardsai.net on Medium, where people are continuing the conversation by highlighting and responding to this story.