{"slug": "i-built-my-friend-a-study-buddy-that-reads-her-notes-and-never-uploads-them", "title": "I Built My Friend a Study Buddy That Reads Her Notes and Never Uploads Them", "summary": "A developer built StudyBuddy, a local-first study assistant for a friend preparing for engineering exams, which ingests lecture slides, scanned handwritten notes and typed summaries into PostgreSQL with pgvector and full-text search so the notes never leave her laptop. The system uses Gemma via Ollama only as a last resort, relying on BGE-M3 embeddings, a bge-reranker-v2-m3 reranker, a semantic cache, code-graded MCQs and Temporal-backed durable ingestion, and the developer reports that switching to small-to-big retrieval fixed an early failure where a \"what is a pointer?\" query returned a wall of unrelated text.", "body_md": "*This is a submission for the [Hacktoberfest Weekend Challenge: Build for a Friend](https://dev.to/challenges/hacktoberfest-weekend-2026-10-01)*\n\nI built this for Kanishka, who is preparing for Engineering 1st sem exams. She has a problem every serious student knows: hundreds of pages of lecture slides, scanned handwritten notes and typed summaries, scattered across folders, none of it searchable. Before every revision session she spends the first hour just *finding* things.\n\nThe obvious fix is to upload everything to a chatbot. She won't, and I agree with her. Her notes hold her draft answers, her mnemonics and her margin scribbles about what she doesn't understand. That is a record of how she thinks, and she shouldn't have to hand it to a server she has never heard of.\n\n**StudyBuddy** is a local-first study assistant:\n\nIt runs on a laptop. After the one-time model download it works with the Wi-Fi off, including on her commute.\n\n```\n{% embed https://drive.google.com/file/d/1YBFgXPnGcMYXTPvZrxzWmpFqQgIdUfEZ/view?usp=sharing %}\n```\n\nThe recording shows:\n\n`/uploads`, with its status flipping from `pending` to `ready`\nHer real notes never leave her machine.\n\n**[Embed your GitHub repo]**\n\n```\n{% embed https://github.com/livanshmalhotra/StudyBuddy.git %}\ndrop file → parse → chunk → embed → pgvector\nquestion  → cache → hybrid search → rerank → gate → (LLM only if needed)\nquiz      → cached questions → code-graded MCQs → weak-topic ranking\n```\n\n**The principle: the LLM is the last resort.** A small local model is slow and sometimes unreliable, so every step that *can* work without it does. Ingestion has no LLM. Search has no LLM. Repeat questions hit a semantic cache. MCQs are graded in plain code. The model only runs when a question genuinely needs synthesis.\n\n| Layer | Tool | Why | \n|---|---|---|\n| LLM | **Gemma** via Ollama | Answers, quiz generation, free-text grading | \n| Embeddings | **BGE-M3** | Multilingual, self-hosted | \n| Reranker | **bge-reranker-v2-m3** | Open-weight, CPU-friendly, calibrated scores | \n| Database | **PostgreSQL + pgvector + full-text search** | Hybrid retrieval in one place | \n| Durable ingestion | **Temporal** | A 600-page scan that fails at page 400 resumes instead of restarting | \n| Backend / UI | **FastAPI** ,**React + Vite** | No lock-in | \n| Tracing | **[Sentry, if your DSN is live]** | Every request tagged `llm_used` | \n\nA watcher monitors the uploads folder. Each file is hashed, so duplicates are skipped and an edited file is re-ingested alone. Nothing is retrained, because adding a document to a RAG system is just an index update.\n\nMy first version answered *\"what is a pointer?\"* with this:\n\nC++ - Quick Notes Page 1 of 2 C++ Syntax, memory, OOP, STL, templates and modern C++ 1. Program Structure & Basics #include using namespace std; int main() { int x = 5; ... 2. Pointers, References & Memory * Pointer stores an address... 3. Classes & OOP class Animal { protected: ...\n\nThe answer was in there, buried in a wall of unrelated text. Four problems were stacking up:\n\nThe fix was **small-to-big retrieval**:\n\nThe same question now returns:\n\n• Pointer stores an address: `int* p = &x;` and `*p` dereferences it. [C++ Quick Notes, p.1]\n\n• Prefer smart pointers over raw new/delete to avoid leaks. [C++ Quick Notes, p.1]\n\nI wrote **[N]** question and expected-answer pairs from her real notes and ran them before and after.\n\n|  | Before | After | \n|---|---|---|\n| Avg. characters shown | **[ ]** | **[ ]** | \n| Answer contains the fact | **[ ]%** | **[ ]%** | \n| Contains unrelated text | **[ ]%** | **[ ]%** | \n| Answered without the LLM | **[ ]%** | **[ ]%** | \n| p50 latency (no LLM / LLM) | **[ ]s / [ ]s** | **[ ]s / [ ]s** | \n\n**[One honest sentence on any number that disappointed you.]**\n\n**Her notes stay hers.** With a closed API, every page of her notes is a request to someone else's server. With open-weight models and a local database, the whole system runs on a laptop with the network unplugged. That is not a feature you can bolt onto a closed model, however good its privacy policy is.\n\n**Zero cost per question.** In exam season she may ask thousands of questions. Free local inference means she never rations curiosity because of a bill.\n\n**I could fix the retrieval because I owned it.** The pointer-dump bug lived in chunking, scoring and the gate. Every layer was readable code I could change. With a black-box RAG API, the best I could have done is tweak a prompt and hope.\n\n**Swappable models.** The model is one line in `.env`. I started on `gemma:2b` and moved to `gemma3:4b` for grounded answering. **[Say what actually changed.]**\n\n**Where a closed model would have been better:** a frontier model writes more fluent explanations and handles messy questions more gracefully than a 2B or 4B model on a laptop CPU. Mine is also slower, about **[ ]s** per synthesized answer. I accepted that because most queries never reach the model, and because the alternative was her notes leaving the device.\n\n**[Link or embed your DevRelay session, and say what it built: e.g. the ingestion workflow, the reranker fix, the gate calibration.]**\n\n**[Write this from real life.]** What was the first question she asked? What broke or surprised her? Which topic did the weak-topics list flag that she hadn't realized was weak? What did she say at the end, word for word?\n\n**Best Use of Temporal:** durable ingestion with retries and resume\n\n**Best Use of Sentry Agent Tracing:** `llm_used` tags and per-stage latency", "url": "https://wpnews.pro/news/i-built-my-friend-a-study-buddy-that-reads-her-notes-and-never-uploads-them", "canonical_source": "https://dev.to/livansh_malhotra/i-built-my-friend-a-study-buddy-that-reads-her-notes-and-never-uploads-them-12di", "published_at": "2026-10-04 18:35:45+00:00", "updated_at": "2026-10-04 18:43:02.086159+00:00", "lang": "en", "topics": ["ai-tools", "ai-products", "mlops", "ai-infrastructure", "large-language-models"], "entities": ["StudyBuddy", "Kanishka", "Gemma", "Ollama", "BGE-M3", "bge-reranker-v2-m3", "PostgreSQL", "pgvector"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-built-my-friend-a-study-buddy-that-reads-her-notes-and-never-uploads-them", "markdown": "https://wpnews.pro/news/i-built-my-friend-a-study-buddy-that-reads-her-notes-and-never-uploads-them.md", "text": "https://wpnews.pro/news/i-built-my-friend-a-study-buddy-that-reads-her-notes-and-never-uploads-them.txt", "jsonld": "https://wpnews.pro/news/i-built-my-friend-a-study-buddy-that-reads-her-notes-and-never-uploads-them.jsonld"}}