LLM-as-Judge in Education: A Curriculum-Grounded Marking Pipeline

wpnews.pro

cd /news/large-language-models/llm-as-judge-in-education-a-curricul… · home › topics › large-language-models › article

[ARTICLE · art-30494] src=arxiv.org ↗ pub=2026-06-17T04:00Z topic=large-language-models verified=true sentiment=↑ positive

LLM-as-Judge in Education: A Curriculum-Grounded Marking Pipeline

Researchers have developed a curriculum-grounded LLM-as-Judge pipeline for automated marking of exam questions, co-developed with an industrial partner. The system uses authorized curriculum artefacts and marking guidelines to generate rubrics and evaluate student responses, achieving results comparable to human tutors. It has been integrated into an online study platform for university admission exam preparation.

read1 min views1 publishedJun 17, 2026

arXiv:2606.17507v1 Announce Type: new Abstract: Generative AI and large language models (LLMs) are increasingly applied to question generation and automated assessment. However, deploying LLMs in preparation for high-stakes exams requires more than prompt engineering; it demands software pipelines that systematically ground model outputs in authorised curriculum artefacts and marking guidelines issued by education authorities. This paper presents a curriculum-grounded, configurable LLM-as-Judge pipeline for question-level marking, co-developed with an industrial partner, to support exam preparation for university admission. The pipeline identifies the relevant topics, subtopics, and cognitive demand of a question, and assembles verifiable and authorised context to support LLM judgement. Curriculum intent is operationalised through concrete syllabus artefacts, including prescribed verbs and outcomes, performance band descriptors, glossary definitions, and marking-guideline principles. A staged LLM workflow is employed to first generate question-specific rubrics, capturing structured expectations of performance, and then derive and evaluate marking criteria used to allocate marks to student responses. This design improves consistency, transparency, and alignment with official marking practices. Preliminary evaluation shows that the proposed LLM-as-Judge pipeline delivers marking outcomes comparable to human tutors, while yielding justifications that are more traceable to authorised curriculum artefacts and marking standards. The pipeline has also been integrated into an online study platform, where early deployment data provide initial insights into operational usage and manual overrides.

source & further reading

arxiv.org — original article

~/api · this article 200

$curl api.wpnews.pro/v1/news/llm-as-judge-in-educatio…

Read original on arxiv.org → arxiv.org/abs/2606.17507

mentioned entities

arXiv

LLM-as-Judge

metadata

slugllm-as-judge-in-education-a-curriculum-grounded-marking-pipeline

topic#large-language-models

secondary3 topics

sentimentpositive

canonicalarxiv.org

navigation

← prevRay Data LLM enables 2x throughp…

next →Trust Begins with DNS: Mitigatin…

── more in #large-language-models 4 stories · sorted by recency

arxiv.org · 17 Jun · #large-language-models

Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation

arxiv.org · 17 Jun · #large-language-models

Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems

arxiv.org · 17 Jun · #large-language-models

Surrogate Assisted Pedestrian Protection Design via a Foundation Model Orchestrated Workflow

arxiv.org · 17 Jun · #large-language-models

DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models

── more on @arxiv 3 stories trending now

wpnews · 16 Jun · #ai-agents

The LLM Is Not the Final Authority: Building Trust Infrastructure for AI Agents

wpnews · 16 Jun · #artificial-intelligence

Most Businesses Lose Leads at Night — So I Built This

wpnews · 16 Jun · #ai-safety

Researchers propose causal framework to audit synthetic data

sponsored brought to you by zahid.host 4,200+ EU-deployed projects

reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main

→ Live at https://your-agent.zahid.host ✓

Get free account → Pricing

from €0/mo · no card required