I Built My Friend a Study Buddy That Reads Her Notes and Never Uploads Them A developer built StudyBuddy, a local-first study assistant for a friend preparing for engineering exams, which ingests lecture slides, scanned handwritten notes and typed summaries into PostgreSQL with pgvector and full-text search so the notes never leave her laptop. The system uses Gemma via Ollama only as a last resort, relying on BGE-M3 embeddings, a bge-reranker-v2-m3 reranker, a semantic cache, code-graded MCQs and Temporal-backed durable ingestion, and the developer reports that switching to small-to-big retrieval fixed an early failure where a "what is a pointer?" query returned a wall of unrelated text. This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend https://dev.to/challenges/hacktoberfest-weekend-2026-10-01 I built this for Kanishka, who is preparing for Engineering 1st sem exams. She has a problem every serious student knows: hundreds of pages of lecture slides, scanned handwritten notes and typed summaries, scattered across folders, none of it searchable. Before every revision session she spends the first hour just finding things. The obvious fix is to upload everything to a chatbot. She won't, and I agree with her. Her notes hold her draft answers, her mnemonics and her margin scribbles about what she doesn't understand. That is a record of how she thinks, and she shouldn't have to hand it to a server she has never heard of. StudyBuddy is a local-first study assistant: It runs on a laptop. After the one-time model download it works with the Wi-Fi off, including on her commute. {% embed https://drive.google.com/file/d/1YBFgXPnGcMYXTPvZrxzWmpFqQgIdUfEZ/view?usp=sharing %} The recording shows: /uploads , with its status flipping from pending to ready Her real notes never leave her machine. Embed your GitHub repo {% embed https://github.com/livanshmalhotra/StudyBuddy.git %} drop file → parse → chunk → embed → pgvector question → cache → hybrid search → rerank → gate → LLM only if needed quiz → cached questions → code-graded MCQs → weak-topic ranking The principle: the LLM is the last resort. A small local model is slow and sometimes unreliable, so every step that can work without it does. Ingestion has no LLM. Search has no LLM. Repeat questions hit a semantic cache. MCQs are graded in plain code. The model only runs when a question genuinely needs synthesis. | Layer | Tool | Why | |---|---|---| | LLM | Gemma via Ollama | Answers, quiz generation, free-text grading | | Embeddings | BGE-M3 | Multilingual, self-hosted | | Reranker | bge-reranker-v2-m3 | Open-weight, CPU-friendly, calibrated scores | | Database | PostgreSQL + pgvector + full-text search | Hybrid retrieval in one place | | Durable ingestion | Temporal | A 600-page scan that fails at page 400 resumes instead of restarting | | Backend / UI | FastAPI , React + Vite | No lock-in | | Tracing | Sentry, if your DSN is live | Every request tagged llm used | A watcher monitors the uploads folder. Each file is hashed, so duplicates are skipped and an edited file is re-ingested alone. Nothing is retrained, because adding a document to a RAG system is just an index update. My first version answered "what is a pointer?" with this: C++ - Quick Notes Page 1 of 2 C++ Syntax, memory, OOP, STL, templates and modern C++ 1. Program Structure & Basics include using namespace std; int main { int x = 5; ... 2. Pointers, References & Memory Pointer stores an address... 3. Classes & OOP class Animal { protected: ... The answer was in there, buried in a wall of unrelated text. Four problems were stacking up: The fix was small-to-big retrieval : The same question now returns: • Pointer stores an address: int p = &x; and p dereferences it. C++ Quick Notes, p.1 • Prefer smart pointers over raw new/delete to avoid leaks. C++ Quick Notes, p.1 I wrote N question and expected-answer pairs from her real notes and ran them before and after. | | Before | After | |---|---|---| | Avg. characters shown | | | | Answer contains the fact | % | % | | Contains unrelated text | % | % | | Answered without the LLM | % | % | | p50 latency no LLM / LLM | s / s | s / s | One honest sentence on any number that disappointed you. Her notes stay hers. With a closed API, every page of her notes is a request to someone else's server. With open-weight models and a local database, the whole system runs on a laptop with the network unplugged. That is not a feature you can bolt onto a closed model, however good its privacy policy is. Zero cost per question. In exam season she may ask thousands of questions. Free local inference means she never rations curiosity because of a bill. I could fix the retrieval because I owned it. The pointer-dump bug lived in chunking, scoring and the gate. Every layer was readable code I could change. With a black-box RAG API, the best I could have done is tweak a prompt and hope. Swappable models. The model is one line in .env . I started on gemma:2b and moved to gemma3:4b for grounded answering. Say what actually changed. Where a closed model would have been better: a frontier model writes more fluent explanations and handles messy questions more gracefully than a 2B or 4B model on a laptop CPU. Mine is also slower, about s per synthesized answer. I accepted that because most queries never reach the model, and because the alternative was her notes leaving the device. Link or embed your DevRelay session, and say what it built: e.g. the ingestion workflow, the reranker fix, the gate calibration. Write this from real life. What was the first question she asked? What broke or surprised her? Which topic did the weak-topics list flag that she hadn't realized was weak? What did she say at the end, word for word? Best Use of Temporal: durable ingestion with retries and resume Best Use of Sentry Agent Tracing: llm used tags and per-stage latency