Automated Assessment of Medical Students’ Competency-Based Performance Using Natural Language Processing (NLP)
Bibliographic Data
Assessment systems in competency-based medical education increasingly rely on narrative feedback to describe learner performance. 1 When collected over time, learner portfolios include hundreds of narrative comments describing behavioral performance such as communication and teamwork that are not evident in categorical ratings or quantitative assessments alone. However, this large volume of data makes summative competency assessment both time and resource intensive, limiting its frequency. 2 Since 2014, our medical school has implemented a summative competency review at the end of M2 year to identify students who could benefit from additional skills training before the clerkship phase of the curriculum. 3 However, this timing may miss the opportunity to provide students support when they have documented areas for growth earlier in the curriculum. This study uses automated methods analyzing both quantitative ratings and narrative feedback to develop an early detection model of concerning performance after M1 year. To our knowledge, this study is the first application of NLP methods to predict performance using longitudinal student assessment data. To train the model, we used 919 M2 students previously reviewed by trained faculty in 2014–2019; to test the model, we selected 319 students previously reviewed in 2020–2021. Reviewers assessed whether students met standards in communication, professionalism, patient care, and teamwork competencies. The binary outcome of the review, whether or not the student met competency standards, served as the dependent variable. In 2014–2019, 21% of students did not meet standards. Two types of features were generated from the data and used in the model: The number of below-benchmark categorical ratings in each competency. Counts of terms associated with an increased or decreased likelihood of meeting all competency standards, aggregated into several topics, e.g., positive adjectives, late/absent, hedging words, positive teamwork. A logistic regression model was used because it is familiar to students and faculty reviewers, allows for uncertainty estimates (standard errors) on predicted values, is computationally easy to run and evaluate, and performed comparably to other algorithms. The early detection model using ratings alone achieved an AUC (area under the ROC curve) of 0.76 on the 2020 test data, while a model incorporating narrative comments achieved an AUC of 0.83. The results were assessed across different groups of students by gender and race/ethnicity to ensure equity. After converting predicted probabilities from the early detection model to binary predictions (likely to meet all standards vs likely to NOT meet all standards), the overall accuracy of the model is 84% for the 2020–2021 test data. A qualitative review found the model identified small numbers of “false negatives”—students whose areas for concern emerged during M2 year and therefore were not detectable by the M1-based early detection model. It also found small numbers of “false positives”—students whose patterns of performance improved from M1 to M2 year and were not determined to need additional support. The model is sufficiently accurate to be useful as an early indicator that a student may benefit from additional support in developing behavioral competencies, though we did identify both false-positive and -negative results during analysis. The application of NLP to student performance data produced a reasonably accurate prediction of students who would benefit from additional skills training after the M1 year. However, while the early detection model will allow faculty to focus efforts on those most likely to need support, we believe human review of the predicted students’ data is still necessary, both to verify the findings and to cater additional training to individual needs. Future research will study the potential emotional impact of “labeling” students as needing support early in the preclerkship curriculum
Categorical variable · Curriculum · Educational measurement · Formative assessment · Machine learning · Mathematics education · Medical education · Narrative · Pedagogy · Summative assessment · Teamwork · Clinical Reasoning and Diagnostic Skills · Computer Science · Innovations in Medical Education · Medical Education and Admissions · Medicine · Psychology
| Citation velocity | historical |
|---|---|
| Highly cited | No |