{"slug": "i-built-an-ai-tutor-that-cannot-do-the-homework", "title": "I Built an AI Tutor That Cannot Do the Homework", "summary": "A developer built a custom AI tutor for his nine-year-old son with ADHD that is architecturally prevented from solving homework: the mentor endpoint receives only task text while correct answers live in a table the LLM layer never reads, and every reply must end with a forward-moving question. The system transcribes tasks verbatim from the child's actual textbook via narrow vision-model prompts with scan-based verification (22 pages, 117 exercises, zero invented wording in six weeks), encodes pedagogy in the task schema through hooks, probes and answer-shaped distractors, and renders model notebook entries on a squared sheet so the child can transfer work to paper for the teacher.", "body_md": "My son is nine, in second grade, and has ADHD. He can calculate, but store-bought trainer apps reward taps, not understanding, and he finishes each one in an evening by guessing. So I built our own tutor around his actual school textbook (the skin is a space school: missions follow the Apollo 11 route, 15 rockets in the collection), and the whole architecture serves a single ban: the LLM is not allowed to solve the task.\n\nThis post is my engineering story: the constraint that shapes everything, the content pipeline, the pedagogy encoded in data, and the surprising part, the paper notebook.\n\n## \n  \n  \n  The ban that shapes the architecture\n\nMost AI homework helpers on the market do the homework. Mine can't, and not because of a polite system prompt. The design is structural, not polite:\n\n- The mentor endpoint receives only the task text. The correct answers live in a table the LLM layer never reads, so there's nothing to leak and nowhere to jailbreak.\n- The prompt must end every reply with one question that moves the child forward. No final answers, no \"the answer is\".\n- The similar-example mode works through an analogous problem with different numbers. It may name that example's answer, never the original task's.\nIf you take one thing from my story: isolate the answer from the model, then design the prompt around a question-ending rule. Bans you can verify in code beat bans you can only write in prose. I trust the first kind.\n## Content pipeline: the textbook is the source of truth\nHomework comes from a specific textbook, so the dataset has to match it verbatim: our numbering, our wording, so the notebook can be shown to the teacher in the morning. I forbid generative paraphrase, and the pipeline enforces it:\npage scan (PNG, 300 dpi)\n-> vision model with a narrow prompt: \"transcribe task N verbatim\"\n-> second pass on disputed fragments\n-> verification against the scan; the scan always wins\n-> a card in the dataset\nFull-page transcription by vision models is unstable. Narrow one-task prompts are reliable, so my pipeline reads pages piece by piece.\nWhy am I so paranoid? Here's a real catch from the first weeks: an external agent prepared five tasks for a page, and verification against the scan flagged one. In an angle-counting task the agent had picked the wrong figure and \"found\" a right angle in a hexagon, while on the scan the right angle belongs to the pentagon. Without that check my son would've memorized a wrong fact and carried it to class.\nSix weeks in: 22 pages across two textbooks, 117 exercises, zero wording invented by feel.\n## Pedagogy encoded in data, not in pixels\nThe method I lean on comes from the open learn method (the amosblomqvist/learn repo): understanding is a connected dependency graph, not a pile of facts. Unconditional truths first, derived steps second, every node motivated before it's established, then linked to the previous one and fixed with a quiz.\nIn my app it lives in the task schema, roughly:\n\nThree fields do the teaching. In practice, the hook motivates (\"inverse problems are a time machine: reconstruct the question from the answer\"). The probe questions come before the hard task: before the cucumber word problem, \"what is 6 + 4?\" My kid with ADHD gets an early win and enters the real task without fear. Distractors are mutations of the correct answer, same length, same shape, each a mistake a child actually makes. When my son picks wrong, he gets an explanation, not a red cross.\n\n## \n  \n  \n  The notebook bridge\n\nThe big discovery of my son's first grade: for half the class the problem isn't arithmetic, it's formatting. The teacher grades the notebook, not the app. So after each solved task the app renders a model entry on a squared sheet: a four-cell margin with a red line, the date on the 11th cell, columns three cells apart, wrapping to sheet width, a handwriting font in blue ink.\n\nMy son solves in the app, where input is easy and feedback is instant, then transfers to paper following the model. Homework is still delivered the way school wants it, and that's non-negotiable.\n\n## \n  \n  \n  Deliberately boring tech\n\nReact 18, TypeScript, Vite, Tailwind. Express with better-sqlite3, cookie sessions, pm2 and nginx. An LLM behind the mentor buttons (Explain, Hint, Similar example; GLM via the z.ai API), replies spoken aloud with speech synthesis. Content is prepared by an agent in the IDE following the pipeline above; a human approves publication. It's the same principle as my grown-up automations: the agent prepares, the human decides.\n\n## \n  \n  \n  Why systems like this are rare\n\nI count five conditions that must hold at once, and the combination is scarce:\n\n1. A teaching method pushed down to data fields, not gamification painted over a PDF.\n2. An agent content pipeline with scan verification, because without it an LLM hallucinates a textbook within a week.\n3. LLM discipline: the mentor is designed around a ban, with answers isolated from the model.\n4. Respect for the school standard: the paper notebook still has to look right.\n5. An engineer-parent, that's me, watching every evening where the system fails.\nThe edtech market is huge, but it's mostly content platforms without an AI loop, in my view, or AI solvers that do the homework for the child. The middle, a machine that teaches how to learn from your own textbook, is nearly empty.\n## What's next\nA live mentor key, new pages as the school program moves on, and the main experiment: the same content model for a whole class, with teachers watching progress through public links.\nIf you're building something like this for your kids, I'd love to trade notes.\n---\nThis article was prepared with AI assistance (drafting and adaptation) and reviewed by the author before publishing.\nOriginally published at [https://automata.sale/en/blog/ai-automation/space-academy-ai-tutor-from-a-textbook/](https://automata.sale/en/blog/ai-automation/space-academy-ai-tutor-from-a-textbook/) tags: aiagents, edtech, llm, react", "url": "https://wpnews.pro/news/i-built-an-ai-tutor-that-cannot-do-the-homework", "canonical_source": "https://dev.to/automatadotsale/i-built-an-ai-tutor-that-cannot-do-the-homework-51ed", "published_at": "2026-10-05 13:43:32+00:00", "updated_at": "2026-10-05 13:48:36.343129+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-tools", "generative-ai", "ai-agents"], "entities": ["amosblomqvist/learn"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-built-an-ai-tutor-that-cannot-do-the-homework", "markdown": "https://wpnews.pro/news/i-built-an-ai-tutor-that-cannot-do-the-homework.md", "text": "https://wpnews.pro/news/i-built-an-ai-tutor-that-cannot-do-the-homework.txt", "jsonld": "https://wpnews.pro/news/i-built-an-ai-tutor-that-cannot-do-the-homework.jsonld"}}