Abstract
Retrieval practice supports learning but requires educators to build large item banks. We compared LLM-generated and human-written retrieval practice items in an introductory psychology course to test whether LLM items match instructor-written ones in quality. LLM items exhibited overall weaker psychometric properties, suggesting that human supervision may remain necessary during item generation for retrieval practice.
- Anthology ID:
- 2026.aimecon-main.21
- Volume:
- [Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers](https://aclanthology.org/volumes/2026.aimecon-main/)
- Month:
- October
- Year:
- 2026
- Address:
- Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
- Editors:
- Joshua Wilson ,Christopher Ormerod ,Magdalen Beiting-Parrish
- Venue:
- [AIME-Con](https://aclanthology.org/venues/aimecon/)
- SIG:
- Publisher:
- National Council on Measurement in Education (NCME)
- Note:
- Pages:
- 194–203
- Language:
- URL:
- [https://aclanthology.org/2026.aimecon-main.21/](https://aclanthology.org/2026.aimecon-main.21/)
- DOI:
- Cite (ACL):
- Marcus Leong, Jennifer Rose, Lisa Dierker, and Antonio Laverghetta Jr.. 2026. A Feasibility Study on Retrieval Practice Using Large Language Models . InProceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers , pages 194–203, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
- Cite (Informal):
- [A Feasibility Study on Retrieval Practice Using Large Language Models](https://aclanthology.org/2026.aimecon-main.21/) (Leong et al., AIME-Con 2026)
- PDF:
- [https://aclanthology.org/2026.aimecon-main.21.pdf](https://aclanthology.org/2026.aimecon-main.21.pdf)