04:00
2026-08-03
arxiv.org
large-language-models
Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs
A study on arXiv (2607.28634v1) found that zero-shot GPT-4.1 with a temperature of 0 achieved the highest item difficulty prediction accuracy among large language models, with a quadratic weighted kapโฆ