StudyBuddy AI: Taming Gemma 2B to Build a Local Quiz Generator for a Friend A developer built StudyBuddy AI, a fully local web app that generates three-question multiple-choice quizzes on any topic using Google's gemma2:2b model via Ollama, a .NET 10 Minimal API backend, and a vanilla JavaScript frontend. To work around the small model's tendency to hallucinate or merge answers when asked for a multi-question JSON array, the backend queries Gemma for one question at a time in a loop and assembles the JSON itself, while option shuffling is handled client-side. Who I Built It For My friend is currently preparing for university exams and technical interviews focusing on databases and C . Reading dry documentation gets boring quickly, so I wanted to build an interactive, multiple-choice quiz partner that tests their knowledge on any given topic. Since it's for a student, it had to be completely free and accessible without internet restrictions. What I Built I built StudyBuddy AI — a lightweight, completely local web application that generates 3-question multiple-choice quizzes on any subject. GitHub Repository: https://github.com/Shadow16Ua/StudyBuddy-AI https://github.com/Shadow16Ua/StudyBuddy-AI The Tech Stack: - AI Model: Google's gemma2:2b running locally via Ollama. - Backend: .NET 10 Minimal API C to handle prompting and JSON parsing. - Frontend: Pure HTML, CSS, and Vanilla JavaScript with a sleek dark mode . 📸 Demo Why Open Innovation Matters Here Choosing an open-weight model like Gemma 2B over a closed API like OpenAI was crucial for this project for three main reasons: 1. Zero Cost & Privacy: My friend can generate hundreds of quizzes without worrying about API limits, subscription fees, or sending their study data to a third-party server. 2. 100% Offline Capability: It runs perfectly on a standard laptop CPU, meaning they can study during commutes or internet outages. 3. The Engineering Challenge Taming the 2B Model : This was the most interesting part. I quickly realized that small 2B models struggle to output complex JSON arrays like 3 questions at once . It would constantly hallucinate structures or merge answers. Because I had full control over the local inference, I was able to completely redesign the backend pipeline. Instead of asking for 3 questions in one prompt, my C backend asynchronously asks Gemma for one question, exactly 3 times in a loop, and then manually constructs a bulletproof JSON array. This guarantees perfect UI rendering every single time. I also shifted the "shuffle options" logic to the frontend to prevent the AI from confusing correct/incorrect indexes. Open source allowed me to iterate rapidly, observe the raw output locally, and engineer a robust wrapper around a lightweight model to make it perform like a much heavier one. Note: I'm submitting this project for the Best Use of Gemma category as well