# Building an AI STEM Solver That Shows Its Work — and Double-Checks It

> Source: <https://dev.to/yashwanthg/building-an-ai-stem-solver-that-shows-its-work-and-double-checks-it-5dj5>
> Published: 2026-10-01 14:34:04+00:00

Ask any AI chatbot a calculus question and you'll get an answer. Maybe even the right one. But a student doesn't need *an* answer — they need to understand *how* to get there, and they need to know the answer is actually correct. Those are two different engineering problems, and I built [Forge](https://gpai-jade.vercel.app) (repo: [Asdfyash1/GPAI](https://github.com/Asdfyash1/GPAI)) to solve both: a full-stack STEM copilot that turns any problem — typed, photographed, or pasted from a URL — into a step-by-step derivation, then has a *second* model independently verify the result.

The stack: **Next.js 16** (App Router, Turbopack), **TypeScript** (strict), **NVIDIA NIM** models (Nemotron, DeepSeek Flash, Llama 3.3) via the Vercel AI SDK, deployed on **Vercel's Hobby tier**. Here's how the solver works, step by step, with the real architecture.

Nobody reads a wall of math. Forge's solver streams the derivation step-by-step over **Server-Sent Events**, so steps appear as they're generated and the UI can reveal them one at a time (collapsible step reveal — "show me the next step" instead of the whole answer).

The streaming endpoint lives at `/api/educate/stream`:

``` js
// src/app/api/educate/stream/route.ts (simplified)
import { requireAuth } from "@/lib/api-guard";
import { streamSolverResponse } from "@/lib/orchestrator";

export async function POST(request: Request) {
  const guard = await requireAuth(request);
  if (!guard.ok) return guard.response;

  const body = await request.json();
  const handle = await streamSolverResponse({ problem: body.problem });
  return new Response(handle.textStream, {
    headers: { "Content-Type": "text/plain; charset=utf-8" },
  });
}
```

On the client, a small `useStream` hook consumes the SSE stream and appends steps to state as they arrive. The response parser (`src/lib/response-parser.ts`) decomposes the LLM's markdown into structured steps using regex — each step becomes a collapsible card in `SolverView.tsx`. That decomposition matters: if the model returns one blob, you can't do progressive reveal, follow-up chips ("Why does this step work?"), or per-step quizzes later.

This is the part most AI solvers skip. Forge runs a **cross-check**: after the primary model produces a solution, a second model independently solves the same problem, and the two are compared. The UI shows one of three verdicts: **agree**, **minor difference**, or **disagree**.

Why it matters: a single model's confident wrong answer is the most dangerous output in education. Two models arriving at the same answer independently is dramatically stronger evidence than one model's confidence score. The cross-check model is configurable via env var:

```
NVIDIA_SOLVER_MODEL=meta/llama-3.3-70b-instruct
NVIDIA_VERIFIER_MODEL=meta/llama-3.3-70b-instruct
```

You can point the verifier at a completely different model family (or even a different provider via the `ADDITIONAL_OPENAI_COMPATIBLE_*` vars) so the check isn't just the same weights agreeing with themselves. When the verdict is "disagree," the student sees that upfront instead of copying a wrong derivation into their homework.

A derivation full of `x^2 + 2x + 1 = 0` in monospace is unreadable. Forge renders formulas with **KaTeX** via `rehype-katex` + `remark-math`, layered on `react-markdown` + `remark-gfm`. The `MathMarkdown.tsx` component handles the whole pipeline, so model output containing `$...$` and `$$...$$` becomes properly typeset math inline.

Small detail, big UX difference: students trust a solution that *looks* like their textbook. Plain-text math looks like a chatbot guessing; typeset math looks like a worked example.

Half of real homework exists as a photo of a notebook or a textbook page. Forge's OCR pipeline:

`MAX_INLINE_IMAGE_BYTES`). Do this in the browser, not the server; it saves bandwidth and avoids gateway rejections.
PDFs get the same treatment via `unpdf` — a pure-JS parser with no native binaries, which matters because Vercel serverless functions don't love native modules.

Here's a trick more AI apps should steal: Forge runs **without any API keys at all**. If `NVIDIA_API_KEY` isn't set, the app falls back to `demo-solver.ts` — deterministic, hand-written demo output that exercises the entire UI (step reveal, KaTeX, quiz, cross-check display) with no model calls.

Why bother? Three reasons: (1) anyone can clone and run it in 30 seconds, which is gold for a portfolio project; (2) frontend development doesn't burn API credits; (3) the demo doubles as a UI test fixture. The env table is honest about what's needed for what:

| Variable | Needed for | 
|---|---|
| `NVIDIA_API_KEY` | All real AI features | 
| `JWT_SECRET` | Auth (self-generated) | 
| `RESEND_API_KEY` | OTP emails (lazy-initialized — build succeeds without it) | 
| `TELEGRAM_BOT_TOKENS` | Cloud storage backend | 

Every AI endpoint sits behind a `requireAuth()` guard that runs four checks in order: payload size (→ 413), origin validation for CSRF (→ 403), JWT cookie verification (→ 401), and per-user rate limiting (→ 429). The API key lives **only** in Vercel env vars — the frontend calls your own `/api/*` routes with an HttpOnly `SameSite=Lax` cookie and never sees a secret.

The rate limits are tuned per endpoint: 30 req/min for most AI routes, 10/min for debate mode (which fires 3 model calls per request). In-memory sliding-window counters per Vercel isolate, pruned periodically.

A few things I'd do differently or that bit me:

`node_modules`, not your memory, when something behaves oddly.` mistralai/mistral-large-3-675b-instruct-2512` isn't in the NVIDIA NIM catalog — every call 404s. The README warns about it because I learned the hard way.
The pattern generalizes beyond homework: **stream structured output, verify with an independent model, render it like the domain expects, and degrade gracefully without keys.** Most AI wrappers stop at "call the API and print the text." The difference between a demo and a tool is everything around the model call — and that's where the interesting engineering lives.

The repo ([Asdfyash1/GPAI](https://github.com/Asdfyash1/GPAI)) has the full source, and the live app is at [gpai-jade.vercel.app](https://gpai-jade.vercel.app) — try photographing a math problem and watch the cross-check verdict.
