Field-Testing Multiple-Choice Questions With AI Examinees: English Grammar Items

Maeda, Hotaka

doi:10.1177/00131644241281053

Field-Testing Multiple-Choice Questions With AI Examinees: English Grammar Items

Article Status

Published

Author/contributor

Maeda, Hotaka (Author)

Title

Field-Testing Multiple-Choice Questions With AI Examinees: English Grammar Items

Abstract

Field-testing is an essential yet often resource-intensive step in the development of high-quality educational assessments. I introduce an innovative method for field-testing newly written exam items by substituting human examinees with artificially intelligent (AI) examinees. The proposed approach is demonstrated using 466 four-option multiple-choice English grammar questions. Pre-trained transformer language models are fine-tuned based on the 2-parameter logistic (2PL) item response model to respond like human test-takers. Each AI examinee is associated with a latent ability θ, and the item text is used to predict response selection probabilities for each of the four response options. For the best modeling approach identified, the overall correlation between the true and predicted 2PL correct response probabilities was .82 (bias = 0.00, root mean squared error = 0.18). The study results were promising, showing that item response data generated from AI can be used to calculate item proportion correct, item discrimination, conduct item calibration with anchors, distractor analysis, dimensionality analysis, and latent trait scoring. However, the proposed approach did not achieve the level of accuracy obtainable with human examinee response data. If further refined, potential resource savings in transitioning from human to AI field-testing could be enormous. AI could shorten the field-testing timeline, prevent examinees from seeing low-quality field-test items in real exams, shorten test lengths, eliminate test security, item exposure, and sample size concerns, reduce overall cost, and help expand the item bank. Example Python code from this study is available on Github: https://github.com/hotakamaeda/ai_field_testing1

Publication

Educational and Psychological Measurement

Date

2024-10-3

Volume

85

Issue

2

Pages

221-244

Journal Abbr

Educ. Psychol. Meas.

DOI

10.1177/00131644241281053

Citation Key

maeda2025a

URL

https://journals.sagepub.com/doi/10.1177/00131644241281053

Accessed

20/05/2025, 20:10

ISSN

0013-1644

Short Title

Field-Testing Multiple-Choice Questions With AI Examinees

Language

en

Library Catalogue

DOI.org (Crossref)

Extra

<标题>: 使用人工智能考生进行选择题实地测试：英语语法题目 <AI Smry>: An innovative method for field-testing newly written exam items by substituting human examinees with artificially intelligent (AI) examinees is introduced, showing that item response data generated from AI can be used to calculate item proportion correct, item discrimination, conduct item calibration with anchors, distractor analysis, dimensionality analysis, and latent trait scoring. Read_Status: New Read_Status_Date: 2026-01-26T11:33:53.529Z

Citation

Maeda, H. (2024). Field-Testing Multiple-Choice Questions With AI Examinees: English Grammar Items. Educational and Psychological Measurement, 85(2), 221–244. https://doi.org/10.1177/00131644241281053

Link to this record

https://aievidencehub.org/lib/SSIYJR7E