This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
A few weeks ago my friend spent weeks preparing for a software engineering interview. They worked really hard, and then didn't get it. The thing that hurt most wasn't the "no". It was that nobody told them what went wrong. Interview rejections come with no feedback, so you're left replaying a 45-minute performance in your head.
I built EdgeMate for them: a private, offline mock interviewer that runs on their laptop. It runs a real interview (understand the problem, discuss an approach, code, debug, discuss complexity, get a debrief) and it never makes them feel stupid for being stuck.
Three things I cared about:
Local Open-Source AI Coding Interviewer
"The code runner checks whether the solution works. The open-source AI handles the human part of the interview."
flowchart TD
subgraph Frontend["Frontend Layer (Dark-First Developer Theme)"]
UI["Web SPA / UI (FastAPI & Uvicorn / Gradio)\n• Progressive Stage Panels\n• Live Python Editor & Examples Drawer\n• Problem Catalog & Progress Dashboard"]
end
subgraph API["Backend & Orchestration Layer"]
ROUTER["REST API Routes (/api/sessions/*)"]
STATE["Session State Machine (InterviewSession)\n• State: UNDERSTANDING ➔ APPROACH ➔ REVIEW ➔ CODING ➔ TESTING ➔ DEBUGGING ➔ COMPLEXITY ➔ DEBRIEF\n• Real-Time Attempt Tracking & Timestamps"]
AGENT["Interviewer Agent (InterviewerAgent)\n• Context-Bounded Prompt Builder\n• Active Persona Adaptation\n• Socratic Guidance (Zero Solution Leaks)"]
end
subgraph LocalAI["Local Open-Source LLM (Ollama)"]
OLLAMA["Ollama Engine (Local CPU/GPU Inference)\n• Default: qwen2.5-coder:3b (Q4_K_M)\n• Alternatives: llama3.2:3b, mistral:7b\n• 100% Offline & Private"]
end
subgraph Execution["Deterministic Execution Engine (Zero LLM Hallucinations)"]
KB_["KB (interview-kb/.py)\n• 119 Problems • 20 DSA Patterns\n• Pre-Materialized Hidden Tests & Large Generators"]
RUNNER["Test Runner
… The one design rule: code checks, the model talks.
My first instinct was to ask an LLM "is this code correct?". That's exactly where small models are worst. They can't reliably trace pointer loops, big inputs, or off-by-one bugs, and they love to dump the full solution. So I split the job:
Candidate code ──► deterministic subprocess runner ──► pass/fail + a failure tag
│
Candidate chat ◄── local LLM (Ollama) ◄── pre-written follow-up question for that tag
tests_passed, tests_total and a failing_tag such as empty, duplicates or large.
The interview itself is a state machine with 7 stages: understanding → approach → approach review → coding → Socratic debugging → complexity → 4-pillar debrief. The editor and the optimal Big-O stay hidden until the right stage, like a real interview. There's also a 3-level hint ladder (question → nudge → concrete invariant).
The knowledge base is 119 problems across 20 patterns, stored as plain JSON. Each has a reference solution, a brute-force solution, tagged tests, hints, and planted "common bug" solutions with the follow-up question for each. I checked it two ways:
That second check caught my own mistakes: a "bug" that turned out to be a valid alternative algorithm, and bugs my tests couldn't actually catch. I fixed the tests and re-verified.
Stack: Python, FastAPI, SQLite for session history and progress, Ollama for local inference, plain files for personas and problems. Personas (Friendly / Standard / Silent) are markdown files with word budgets.
1. Privacy where it matters most.
Someone processing a rejection is typing "I froze on this question" and their half-working code into a tool. With a hosted API, all of that goes to a server they don't control. With Ollama, nothing leaves the laptop. There's no account, no telemetry, and it works with no internet.
2. It runs on the hardware my friend actually has.
They have an 8GB laptop, so I couldn't pick a big model. The architecture makes that fine: deterministic tools do the judging and a 3B open model does the talking.
3. I can change how the interviewer behaves.
The interviewer's personality is a markdown file. I can make it warmer, stricter, or nearly silent for stress practice, by editing a text file. The problem bank is the same: when my friend gets asked something new next time, they add one JSON entry. No retraining, no subscription.
4. I could compare models instead of trusting one.
EdgeMate has a benchmark endpoint that compares two local models on latency and on whether they respect the tone and word limits.
5. Zero cost to practise.
Fifty practice rounds cost the same as one: nothing. For someone job hunting, that matters.
Where open beat closed, and where it didn't. A hosted frontier model would write a smoother conversation. But it would also cost money per session, send a vulnerable person's code and feelings to a third party, and tempt me to let it judge correctness, which is the part models get wrong. Open weights plus deterministic tools gave me something more trustworthy, more private, and cheaper for this one person.