Building an AI STEM Solver That Shows Its Work — and Double-Checks It A developer built Forge, an open-source full-stack STEM copilot that streams step-by-step derivations over Server-Sent Events and runs a second, independently configurable model as a cross-check, surfacing an agree/minor-difference/disagree verdict to flag confident wrong answers. The Next.js 16 and TypeScript app uses NVIDIA NIM models (Nemotron, DeepSeek Flash, Llama 3.3) via the Vercel AI SDK, renders math with KaTeX, and handles photographed or PDF homework through a browser-side OCR pipeline. The verifier model is set through environment variables so it can be pointed at a different model family or provider. Ask any AI chatbot a calculus question and you'll get an answer. Maybe even the right one. But a student doesn't need an answer — they need to understand how to get there, and they need to know the answer is actually correct. Those are two different engineering problems, and I built Forge https://gpai-jade.vercel.app repo: Asdfyash1/GPAI https://github.com/Asdfyash1/GPAI to solve both: a full-stack STEM copilot that turns any problem — typed, photographed, or pasted from a URL — into a step-by-step derivation, then has a second model independently verify the result. The stack: Next.js 16 App Router, Turbopack , TypeScript strict , NVIDIA NIM models Nemotron, DeepSeek Flash, Llama 3.3 via the Vercel AI SDK, deployed on Vercel's Hobby tier . Here's how the solver works, step by step, with the real architecture. Nobody reads a wall of math. Forge's solver streams the derivation step-by-step over Server-Sent Events , so steps appear as they're generated and the UI can reveal them one at a time collapsible step reveal — "show me the next step" instead of the whole answer . The streaming endpoint lives at /api/educate/stream : js // src/app/api/educate/stream/route.ts simplified import { requireAuth } from "@/lib/api-guard"; import { streamSolverResponse } from "@/lib/orchestrator"; export async function POST request: Request { const guard = await requireAuth request ; if guard.ok return guard.response; const body = await request.json ; const handle = await streamSolverResponse { problem: body.problem } ; return new Response handle.textStream, { headers: { "Content-Type": "text/plain; charset=utf-8" }, } ; } On the client, a small useStream hook consumes the SSE stream and appends steps to state as they arrive. The response parser src/lib/response-parser.ts decomposes the LLM's markdown into structured steps using regex — each step becomes a collapsible card in SolverView.tsx . That decomposition matters: if the model returns one blob, you can't do progressive reveal, follow-up chips "Why does this step work?" , or per-step quizzes later. This is the part most AI solvers skip. Forge runs a cross-check : after the primary model produces a solution, a second model independently solves the same problem, and the two are compared. The UI shows one of three verdicts: agree , minor difference , or disagree . Why it matters: a single model's confident wrong answer is the most dangerous output in education. Two models arriving at the same answer independently is dramatically stronger evidence than one model's confidence score. The cross-check model is configurable via env var: NVIDIA SOLVER MODEL=meta/llama-3.3-70b-instruct NVIDIA VERIFIER MODEL=meta/llama-3.3-70b-instruct You can point the verifier at a completely different model family or even a different provider via the ADDITIONAL OPENAI COMPATIBLE vars so the check isn't just the same weights agreeing with themselves. When the verdict is "disagree," the student sees that upfront instead of copying a wrong derivation into their homework. A derivation full of x^2 + 2x + 1 = 0 in monospace is unreadable. Forge renders formulas with KaTeX via rehype-katex + remark-math , layered on react-markdown + remark-gfm . The MathMarkdown.tsx component handles the whole pipeline, so model output containing $...$ and $$...$$ becomes properly typeset math inline. Small detail, big UX difference: students trust a solution that looks like their textbook. Plain-text math looks like a chatbot guessing; typeset math looks like a worked example. Half of real homework exists as a photo of a notebook or a textbook page. Forge's OCR pipeline: MAX INLINE IMAGE BYTES . Do this in the browser, not the server; it saves bandwidth and avoids gateway rejections. PDFs get the same treatment via unpdf — a pure-JS parser with no native binaries, which matters because Vercel serverless functions don't love native modules. Here's a trick more AI apps should steal: Forge runs without any API keys at all . If NVIDIA API KEY isn't set, the app falls back to demo-solver.ts — deterministic, hand-written demo output that exercises the entire UI step reveal, KaTeX, quiz, cross-check display with no model calls. Why bother? Three reasons: 1 anyone can clone and run it in 30 seconds, which is gold for a portfolio project; 2 frontend development doesn't burn API credits; 3 the demo doubles as a UI test fixture. The env table is honest about what's needed for what: | Variable | Needed for | |---|---| | NVIDIA API KEY | All real AI features | | JWT SECRET | Auth self-generated | | RESEND API KEY | OTP emails lazy-initialized — build succeeds without it | | TELEGRAM BOT TOKENS | Cloud storage backend | Every AI endpoint sits behind a requireAuth guard that runs four checks in order: payload size → 413 , origin validation for CSRF → 403 , JWT cookie verification → 401 , and per-user rate limiting → 429 . The API key lives only in Vercel env vars — the frontend calls your own /api/ routes with an HttpOnly SameSite=Lax cookie and never sees a secret. The rate limits are tuned per endpoint: 30 req/min for most AI routes, 10/min for debate mode which fires 3 model calls per request . In-memory sliding-window counters per Vercel isolate, pruned periodically. A few things I'd do differently or that bit me: node modules , not your memory, when something behaves oddly. mistralai/mistral-large-3-675b-instruct-2512 isn't in the NVIDIA NIM catalog — every call 404s. The README warns about it because I learned the hard way. The pattern generalizes beyond homework: stream structured output, verify with an independent model, render it like the domain expects, and degrade gracefully without keys. Most AI wrappers stop at "call the API and print the text." The difference between a demo and a tool is everything around the model call — and that's where the interesting engineering lives. The repo Asdfyash1/GPAI https://github.com/Asdfyash1/GPAI has the full source, and the live app is at gpai-jade.vercel.app https://gpai-jade.vercel.app — try photographing a math problem and watch the cross-check verdict.