cd /news/artificial-intelligence/part-1-3-from-raw-health-text-to-a-w… Β· home β€Ί topics β€Ί artificial-intelligence β€Ί article
[ARTICLE Β· art-70642] src=dev.to β†— pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Part 1-3: From Raw Health Text to a Working Retrieval Pipeline

A first-year undergraduate student in Artificial Intelligence Engineering is building a Retrieval-Augmented Generation (RAG) project from scratch, documenting the process in a series of blog posts. In the first three parts, the student implemented a pipeline that converts raw health text into structured data, chunks it, generates embeddings, and stores them in a Chroma vector store for retrieval. The project, called health-rag-assistant, is available on GitHub.

read3 min views1 publishedJul 23, 2026

I'm a first-year undergraduate student in Artificial Intelligence Engineering, and this summer I'm building a small RAG (Retrieval-Augmented Generation) project from scratch to learn how these systems actually work under the hood β€” not just by calling an API, but by building the pipeline piece by piece. This is a log of the first three parts of that journey.

#

Part 1: Understanding the Building Blocks

Before writing any real project code, I spent time understanding the core ideas behind RAG:

Calling an LLM API: sending a prompt programmatically and getting a response back. I used the Gemini API for this. llm_test.py

Embeddings: the idea that text can be converted into numerical vectors, where semantically similar sentences end up close to each other in vector space. I tested this with a handful of sentences using sentence-transformers

and computed cosine similarity between them β€” it was satisfying to see related sentences actually cluster together numerically. embedding_test.py

Why not just dump everything into the LLM?: context windows are limited, it's expensive, and irrelevant information can actually hurt answer quality rather than help it. #

Vector databases: conceptually, tools like Chroma or FAISS solve the problem of searching through large numbers of embeddings quickly to find the most relevant ones. No heavy coding in Part 1 β€” mostly small test scripts to confirm I understood each piece before combining them.

#

Part 2: Collecting and Preparing Real Data

With the concepts in place, Part 2 was about getting real data ready for retrieval:

Data collection: I gathered short health topic descriptions (e.g. diabetes, epilepsy) and saved them as structured JSON records, each with a source

and text

field. Keeping the source attached to each record matters β€” it means the system can eventually point back to where an answer came from. data/health_data.json

Chunking: I wrote a script to split each record's text into smaller, paragraph-level pieces while keeping the original source attached to every chunk. Since I intentionally kept the raw text short for this first test run, most entries ended up as a single chunk each β€” which is fine for now, since the goal was to validate the pipeline, not build a large dataset yet. #

Why chunk at all?: embedding models represent short, focused pieces of text more accurately than long documents. Splitting text into meaningful chunks is what makes retrieval actually useful later.

#

Part 3: Setting Up the Vector Store and First Retrieval

With chunked, source-tagged data ready, Part 3 turned that data into something actually searchable:

Vector store setup: installed and configured Chroma, then embedded every chunk from Part 2 and stored it alongside its text and source metadata. embed_and_store.py

First retrieval test: wrote a query script, asked a sample question, and retrieved the most similar chunks from the vector store. The results were reasonable given how small and short the dataset still is β€” a good early sign that the pipeline itself works end to end. retrieval_test.py

A note on limitations: since the dataset is intentionally tiny at this stage, retrieval quality is limited. That's expected, and it'll improve as the dataset grows in later parts.

#

What's Next

At this point I have a full (if small-scale) pipeline: raw text β†’ structured data β†’ chunks β†’ embeddings β†’ retrieval. The next step is connecting retrieval to actual answer generation β€” feeding retrieved chunks into the LLM as context so it can answer health questions grounded in real source material.

πŸ‘‰ Code for this project: health-rag-assistant Follow along as I build this project part by part β€” code on GitHub, progress here on Dev.to.

── more in #artificial-intelligence 4 stories Β· sorted by recency
── more on @gemini 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/part-1-3-from-raw-he…] indexed:0 read:3min 2026-07-23 Β· β€”