cd /news/ai-tools/build-an-art-recommender-from-scratc… · home › topics › ai-tools › article
[ARTICLE · art-144789] src=pub.towardsai.net ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

Build an Art Recommender from scratch with CLIP and ChromaDB

A three-part tutorial series shows developers how to build an art and museum recommender from scratch using CLIP embeddings and ChromaDB, with part one covering data retrieval from the Rijksmuseum API. The author, writing from personal experience as a junior data science practitioner, notes that museums including the Art Institute of Chicago and the Metropolitan Museum of Art have restricted API requests because agents are flooding public endpoints, leaving the Rijksmuseum as a well-documented alternative. The series promises a working laptop prototype, with part two covering vector database construction and part three the adaptive recommender system.

by read9 min views1 publishedOct 4, 2026

A couple of years ago, as I was taking my first steps in the data science world, I had an idea and decided to see if I could build a full-fledged art and museum recommender with no prior knowledge. Yes, none. It was a challenge. I had never set up a database, didn’t know how to collect, what do with the data if I had it, or what were the strategies to provide recommendations. I was starting from scratch. And that is why this project taught me so much. Looking back, it was one of the best learning experiences I’ve had in my (brief) career so far. I believe this kind of project could be a great opportunity for anyone wanting to move from junior ML practitioner to a more experienced engineer, with end-to-end competencies and knowledge of recommender systems. And as a bonus, admire beautiful art along the way. That is why I’ve decided to write this three-part account of the project. If you follow it closely, I can assure you that you will end up with a better sense of how recommender systems work, and you’ll have a working prototype on your laptop.

Since I want to take you through the entire project, this is going to be part one of three. Part two will explore how to turn your images into a machine-readable format and build a vector database with ChromaDB, while part three will put it all together to create an adaptive recommender system.

If this caught your attention, read on!

The first thing I learnt when I decided to start with this project was that art data is ideal: easily available, well-curated, free, and (mostly) copyright-free! What more could a data scientist want? It turns out that several museums provide data from their collections as a service to the general public. So if you want to build something on top of it, you just go and connect to their API. And if you have lots of high-quality data, sky is the limit. Since then things have changed unfortunately. Agents are running wild on the web and any public API is flooded with requests. This is why museums like ARTIC or the MET now restrict requests via their API.

Luckily for us, there are still some museums providing this data via API, and one of them is the Rijksmusem in Amsterdam, which just so happens to be one with a beautiful collection.

So now that we know the data exists, and that the Rijksmuseum API is well-documented, let’s get some data.

After playing a little bit with the Rijksmuseum API, as well as those from other museums, it made sense to build a reference class that contained most of the important logic to download and do a first selection of the artworks to be used for the project. Some of the requirements were:

The GenericMuseumApiInteractor class creates the blueprint for these functionalities. But let’s take a closer look at some of the key points.

def get_response_dict(self, url: str, n_trial: int = 0) -> dict:        json_data = {}        try:            # include a timeout to avoid getting stuck on one request            response = requests.get(url, timeout=10)            # Check if the request was successful            if response.status_code == 200:                json_data: dict = response.json()            else:                print(f"Error: {response.status_code} - {response.text}")        # if the request timed out, try again for a max. of 5 times        except requests.Timeout:            if n_trial < 5:                print(f"Trial {n_trial} failed. Sending request again")                self.get_response_dict(url=url, n_trial=n_trial+1)            else:                print("Unable to reach link")        return json_data

This first function is the one used to get the data from the museum API. Some relevant aspects are:

Another valuable operation is removing duplicates:

def download_image(self, url: str, image_name: str, n_trial: int = 0) -> bool:        try:            response = requests.get(url, timeout=10)            # Check if the request was successful            if response.status_code == 200:                # Get the content of the response                image_data = response.content                if ".jpg" in str(image_name):                    image_name = image_name.rstrip(".jpg")                # Open the image to check for duplicates                img_data = Image.open(io.BytesIO(image_data))                image_hash = imagehash.phash(img_data)                if image_hash not in self.hashes:                    self.hashes.append(image_hash)                    # Specify the local file path where you want to save the image                    local_file_path: str = f"{self.museum_name}/images/{image_name}.jpg"                    # Save the image data to the local file                    with open(local_file_path, "wb") as image_file:                        image_file.write(image_data)         [...]

Using imagehash we check whether the image is identical to one we’ve already downloaded. It’s better to do this as early as possible to avoid polluting your data. No one likes to be recommended the same image five times.

The last segments of this class worth checking are:

def is_exception(self, artwork_dict: dict) -> bool:

and

def passes_requirements(self, artwork_dict: dict, is_nested: bool = False) -> bool:

simply because it’s better to do a bit of curation upfront to remove data that would reduce the quality of your final product. These two functions help with selecting only public domain images, as well as removing, for instance, fragments of artefacts. All of these aspects can be controlled via museum-level configurations. Here is an example of how these can be defined:

museums_details = {    "met": {        "base_url": "https://collectionapi.metmuseum.org/public/collection/v1",        "key": None,        "exceptions": {            "isPublicDomain": False,            "objectName": ["fragment"],            "title": ["fragment"],        },        "requirements": {},        "results_per_page": None,    },    "rijks": {        "base_url": "https://data.rijksmuseum.nl/search/collection?imageAvailable=true",        "key": None,        "exceptions": {},        "requirements": {            "subject_to[0].classified_as[0]._label": "public domain",        },    }}

In general, I would recommend using YAML files to control configurations; they are more light-weight and standardised. Here, I have used a .py file for simplicity. The idea is to create a structure that covers the most important configurations needed to interact with the API of different museums.

Having created this generic blueprint class for interacting with museum APIs, we can move on to create a specific class to interact with the Rijksmuseum API:

from data_retrieval.src.generic_museum_api_interactor import GenericMuseumApiInteractorclass RijksApiInteractor(GenericMuseumApiInteractor):    def __init__(self) -> None:        super().__init__(museum_name="rijks")    def run_down_pipeline(self, n_images_to_download: int) -> None:                super().run_down_pipeline()        page_url = self.base_url        # Loop over the pages        while n_images_to_download > 0:            # structure the query parameters in the correct format            # extract the data for all the images on this page            page_response: dict = self.get_response_dict(url=page_url)            if page_response == {}:                continue            # go through every result in the page            for item in page_response["orderedItems"]:                                artwork_id: str = item["id"].split("/")[-1]                artwork_dict: dict = self.get_response_dict(url=item["id"])                if artwork_dict == {}:                    print(f"Artwork {artwork_id} not found")                    continue                visual_item_dict = self.get_response_dict(                    url=artwork_dict["shows"][0]["id"]                )                if visual_item_dict == {}:                    continue                if not self.passes_requirements(visual_item_dict, is_nested=True):                    continue                digitally_shown_id = visual_item_dict["digitally_shown_by"][0]["id"]                image_dict = self.get_response_dict(url=digitally_shown_id)                image_url = image_dict["access_point"][0]["id"]                if image_url is None or image_url.strip() == "":                    continue                image_downloaded = self.download_image(url=image_url, image_name=artwork_id)                if image_downloaded:                    self.download_json(response_dict=artwork_dict, image_name=artwork_id)                    n_images_to_download -= 1                    if n_images_to_download == 0:                        break            # go to the next page            page_url = page_response["next"]["id"]

The main reason to create specific classes is because every museum has a different design in how their API is structured. In the Rijks one, for instance, results are displayed per page, and only by moving to the next page you get to access more artworks. Elsewhere, you might encounter all results to be displayed as a list of ID’s you need to access one by one.

At this point, we are essentially done! Just go to the root directory of your project (or in a notebook), run the following:

from data_retrieval.src.rijks_api_interactor import RijksApiInteractorinteractor = RijksApiInteractor()interactor.run_down_pipeline(n_images_to_download=100)

and watch the magic happen in front of your eyes. You should see images being downloaded with their corresponding metadata as json files automatically!

Another good option to retrieve images and related data is the Wikimedia Commons API. I haven’t tested it myself, but I’m sure it would add a lot of value after seeing other projects that have used it.

Note: This part is not crucial to be able to proceed further. So if you’d rather dive and learn how to build the core of your recommender engine, feel free to skip directly to the end.

There is one more aspect that is good to address for this kind of project; and that is data consolidation. If you plan to retrieve data from multiple sources (especially public APIs) that data will be structured in different ways. Some APIs might include dates as a list, others as two separate fields. If you want to work across these sources, you need to standardise the data you retrieve. This is made even harder by the fact that these API structures change constantly. So be prepared to refactor your code regularly!

When I was working on this project, I found that this kind of structure worked well for me:

final_metadata_template = {        "original_id": "",        "title_original_language": "",        "title_english": "",        "artists": [],        "main_artist": "",        "artists_details": [],        "museum_api_from_which_it_was_retrieved": "",        "location": "",        "year_start_number": -99999,        "year_start": "",        "year_end_number": -99999,        "year_end": "",        "dating_of_first_display": "",        "production_places": [],        "main_production_place": "",        "dimensions": [],        "artwork_style": [],        "artwork_subjects": [],        "artwork_materials": [],        "artwork_themes": [],        "artwork_techniques": [],        "artwork_types": [],        "artwork_categories": [],        "subject_description_original_language": "",        "short_description_original_language": "",        "long_description_original_language": "",        "other_descriptions_original_language": [],        "subject_description_english": "",        "short_description_english": "",        "long_description_english": "",        "other_descriptions_english": [],        "colours": [],        "normalised_colours": [],        "location_within_museum": "",        "department_within_museum": "",        "date_processed": "",        "documents_and_publications": []        }

In general, select a few fields that you are interested in, such as title, description, author, etc. You can always add more later.

Once you have created your ideal structure, you can create standardised json files for your metadata. I recommend using an LLM to help you with the logic, because looking across multiple json files to create a standard structure, finding the correct mapping of every field etc., is tricky. LLMs are much better than us at these kinds of tasks. Just feed it a couple of the jsons you have downloaded and the final structure you want, and let the tool do the work for you.

Congratulations! You have now laid the foundations on top of which you’re going to build your recommender. I am aware this first part might not have been the most exciting, but your recommender is only as good as the artworks it is built upon. And now, you have high-quality, open-access and curated data that you can use to construct your database. In the next part, we’re going to turn our data into vectors. Let’s get building!

Build an Art Recommender from scratch with CLIP and ChromaDB was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

── more in #ai-tools 4 stories · sorted by recency
── more on @clip 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/build-an-art-recomme…] indexed:0 read:9min 2026-10-04 · —