# I Built My Friend a Study Buddy That Reads Her Notes and Never Uploads Them

> Source: <https://dev.to/livansh_malhotra/i-built-my-friend-a-study-buddy-that-reads-her-notes-and-never-uploads-them-12di>
> Published: 2026-10-04 18:35:45+00:00

*This is a submission for the [Hacktoberfest Weekend Challenge: Build for a Friend](https://dev.to/challenges/hacktoberfest-weekend-2026-10-01)*

I built this for Kanishka, who is preparing for Engineering 1st sem exams. She has a problem every serious student knows: hundreds of pages of lecture slides, scanned handwritten notes and typed summaries, scattered across folders, none of it searchable. Before every revision session she spends the first hour just *finding* things.

The obvious fix is to upload everything to a chatbot. She won't, and I agree with her. Her notes hold her draft answers, her mnemonics and her margin scribbles about what she doesn't understand. That is a record of how she thinks, and she shouldn't have to hand it to a server she has never heard of.

**StudyBuddy** is a local-first study assistant:

It runs on a laptop. After the one-time model download it works with the Wi-Fi off, including on her commute.

```
{% embed https://drive.google.com/file/d/1YBFgXPnGcMYXTPvZrxzWmpFqQgIdUfEZ/view?usp=sharing %}
```

The recording shows:

`/uploads`, with its status flipping from `pending` to `ready`
Her real notes never leave her machine.

**[Embed your GitHub repo]**

```
{% embed https://github.com/livanshmalhotra/StudyBuddy.git %}
drop file → parse → chunk → embed → pgvector
question  → cache → hybrid search → rerank → gate → (LLM only if needed)
quiz      → cached questions → code-graded MCQs → weak-topic ranking
```

**The principle: the LLM is the last resort.** A small local model is slow and sometimes unreliable, so every step that *can* work without it does. Ingestion has no LLM. Search has no LLM. Repeat questions hit a semantic cache. MCQs are graded in plain code. The model only runs when a question genuinely needs synthesis.

| Layer | Tool | Why | 
|---|---|---|
| LLM | **Gemma** via Ollama | Answers, quiz generation, free-text grading | 
| Embeddings | **BGE-M3** | Multilingual, self-hosted | 
| Reranker | **bge-reranker-v2-m3** | Open-weight, CPU-friendly, calibrated scores | 
| Database | **PostgreSQL + pgvector + full-text search** | Hybrid retrieval in one place | 
| Durable ingestion | **Temporal** | A 600-page scan that fails at page 400 resumes instead of restarting | 
| Backend / UI | **FastAPI** ,**React + Vite** | No lock-in | 
| Tracing | **[Sentry, if your DSN is live]** | Every request tagged `llm_used` | 

A watcher monitors the uploads folder. Each file is hashed, so duplicates are skipped and an edited file is re-ingested alone. Nothing is retrained, because adding a document to a RAG system is just an index update.

My first version answered *"what is a pointer?"* with this:

C++ - Quick Notes Page 1 of 2 C++ Syntax, memory, OOP, STL, templates and modern C++ 1. Program Structure & Basics #include using namespace std; int main() { int x = 5; ... 2. Pointers, References & Memory * Pointer stores an address... 3. Classes & OOP class Animal { protected: ...

The answer was in there, buried in a wall of unrelated text. Four problems were stacking up:

The fix was **small-to-big retrieval**:

The same question now returns:

• Pointer stores an address: `int* p = &x;` and `*p` dereferences it. [C++ Quick Notes, p.1]

• Prefer smart pointers over raw new/delete to avoid leaks. [C++ Quick Notes, p.1]

I wrote **[N]** question and expected-answer pairs from her real notes and ran them before and after.

|  | Before | After | 
|---|---|---|
| Avg. characters shown | **[ ]** | **[ ]** | 
| Answer contains the fact | **[ ]%** | **[ ]%** | 
| Contains unrelated text | **[ ]%** | **[ ]%** | 
| Answered without the LLM | **[ ]%** | **[ ]%** | 
| p50 latency (no LLM / LLM) | **[ ]s / [ ]s** | **[ ]s / [ ]s** | 

**[One honest sentence on any number that disappointed you.]**

**Her notes stay hers.** With a closed API, every page of her notes is a request to someone else's server. With open-weight models and a local database, the whole system runs on a laptop with the network unplugged. That is not a feature you can bolt onto a closed model, however good its privacy policy is.

**Zero cost per question.** In exam season she may ask thousands of questions. Free local inference means she never rations curiosity because of a bill.

**I could fix the retrieval because I owned it.** The pointer-dump bug lived in chunking, scoring and the gate. Every layer was readable code I could change. With a black-box RAG API, the best I could have done is tweak a prompt and hope.

**Swappable models.** The model is one line in `.env`. I started on `gemma:2b` and moved to `gemma3:4b` for grounded answering. **[Say what actually changed.]**

**Where a closed model would have been better:** a frontier model writes more fluent explanations and handles messy questions more gracefully than a 2B or 4B model on a laptop CPU. Mine is also slower, about **[ ]s** per synthesized answer. I accepted that because most queries never reach the model, and because the alternative was her notes leaving the device.

**[Link or embed your DevRelay session, and say what it built: e.g. the ingestion workflow, the reranker fix, the gate calibration.]**

**[Write this from real life.]** What was the first question she asked? What broke or surprised her? Which topic did the weak-topics list flag that she hadn't realized was weak? What did she say at the end, word for word?

**Best Use of Temporal:** durable ingestion with retries and resume

**Best Use of Sentry Agent Tracing:** `llm_used` tags and per-stage latency
