cd /news/artificial-intelligence/teachmategpt-a-multi-agent-knowledge… · home topics artificial-intelligence article
[ARTICLE · art-99333] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

TeachMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials

TeachMateGPT, a multi-agent framework introduced in an arXiv paper (arXiv:2608.13708v1), improves curriculum-grounded science assessment generation by raising faithfulness from 0.68 to 0.96 and answer relevancy from 0.60 to 0.89 over a vanilla RAG baseline. The system introduces COPE, a hierarchical knowledge base; a staged fail-closed agent pipeline; SAVER, a source-attributed verification protocol; and NCTB-SciGen8, a dataset of 198 items spanning all 14 chapters of the NCTB Class 8 science textbook, rated by three practicing teachers.

read1 min views1 publishedAug 17, 2026

arXiv:2608.13708v1 Announce Type: new Abstract: Automatically generating textbook-grounded assessment items can reduce science teachers' workload, but existing retrieval-augmented generation (RAG) systems rely on flat retrieval, support only single-question generation, lack safeguards against weak evidence, and are ill-suited to low-resource, board-exam-structured curricula. We address these limitations with TeachMateGPT, a multi-agent system contributing four advances to curriculum-grounded science-assessment authoring. (i) COPE, a hierarchical knowledge base replacing token-window chunking with a multi-resolution index that segments documents along syllabus structure and links them at three granularities via a traversable graph-based lineage, matching evidence to each topic's instructional level. (ii) A staged, fail-closed agent pipeline replacing one-shot retrieve-then-generate: routing gates search, retrieval fuses dense and lexical evidence under a coverage gate that withholds generation on insufficient evidence, and specialist agents draft objective and constructed-response items. (iii) SAVER, a source-attributed verification protocol scoring faithfulness, relevance, and hallucination risk against retrieved evidence, applying stricter grounding checks across each creative question's four sub-parts, paired with teacher-in-the-loop evaluation rather than automatic filtering. (iv) NCTB-SciGen8, a curriculum-grounded dataset of 198 items (143 multiple-choice, 55 creative questions) spanning all 14 chapters of the NCTB Class 8 science textbook, produced by the pipeline and rated by three practicing teachers. TeachMateGPT raises faithfulness (0.68 $\rightarrow$ 0.96) and answer relevancy (0.60 $\rightarrow$ 0.89) over a vanilla RAG baseline.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @teachmategpt 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/teachmategpt-a-multi…] indexed:0 read:1min 2026-08-17 ·