AI-Assisted Exam Variant Generation
A Human-in-the-Loop Framework for Automatic Item Creation
Bibliographic Data
| ID | 22044568 |
|---|---|
| Authors | Charles MacDonald Burke (0009-0003-5510-7809, Franklin University Switzerland, corresponding author) |
| Year | 2025 |
| Volume | 15 |
| Issue | 8 |
| Pages | 1029 |
| Publication date | 2025-08-11 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Education Sciences (JOURNAL) |
| Journal identifiers | ISSN: 2227-7102 • E-ISSN: 2227-7102 |
| Publisher | MDPI AG (PUBLISHER • IT) |
| DOI | 10.3390/educsci15081029 |
| OpenAlex | W4413215135 |
| Language | EN |
| Citations received | 2 |
| References cited | 20 |
Educational assessment relies on well-constructed test items to measure student learning accurately, yet traditional item development is time-consuming and demands specialized psychometric expertise. Automatic item generation (AIG) offers template-based scalability, and recent large language model (LLM) advances promise to democratize item creation. However, fully automated approaches risk introducing factual errors, bias, and uneven difficulty. To address these challenges, we propose and evaluate a hybrid human-in-the-loop (HITL) framework for AIG that combines psychometric rigor with the linguistic flexibility of LLMs. In a Spring 2025 case study at Franklin University Switzerland, the instructor collaborated with ChatGPT (o4-mini-high) to generate parallel exam variants for two undergraduate business courses: Quantitative Reasoning and Data Mining. The instructor began by defining “radical” and “incidental” parameters to guide the model. Through iterative cycles of prompt, review, and refinement, the instructor validated content accuracy, calibrated difficulty, and mitigated bias. All interactions (including prompt templates, AI outputs, and human edits) were systematically documented, creating a transparent audit trail. Our findings demonstrate that a HITL approach to AIG can produce diverse, psychometrically equivalent exam forms with reduced development time, while preserving item validity and fairness, and potentially reducing cheating. This offers a replicable pathway for harnessing LLMs in educational measurement without sacrificing quality, equity, or accountability
Accountability · Audit · Data science · Database · Human-in-the-loop · Scalability · Workflow · Artificial Intelligence in Healthcare and Education · Computer Science · Intelligent Tutoring Systems and Adaptive Learning · Online Learning and Analytics · Artificial Intelligence
Survey of Hallucination in Natural Language Generation
Item response theory.
ChatGPT for good? On opportunities and challenges of large language models for education
Automatic item generation
Closing the Gap
Engineered Prompts in ChatGPT for Educational Assessment in Software Engineering and Computer Science
Automatic item generation for educational assessments
Practical and ethical challenges of large language models in education
| Unique citing works | 2 |
|---|---|
| Citations per year | 2 |
| Citation span | 2026 - 2026 (1) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 2 |