Utilizing large language models for EFL essay grading
An examination of reliability and validity in rubric‐based assessments
Bibliographic Data
| ID | 21297215 |
|---|---|
| Authors | Fatih Yavuz (0000-0003-2645-2710, Preparatory Department Mudanya University Mudanya Turkey, corresponding author), Özgür Çelik (0000-0002-0300-9073, School of Foreign Languages Balıkesir University Balıkesir Turkey), Gamze Yavaş Çelik (0000-0003-1571-9686, School of Foreign Languages Balıkesir University Balıkesir Turkey) |
| Year | 2025 |
| Volume | 56 |
| Issue | 1 |
| Pages | 150-166 |
| Publication date | 2025-01-01 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | British Journal of Educational Technology (JOURNAL) |
| Journal identifiers | ISSN: 0007-1013 • E-ISSN: 1467-8535 |
| Publisher | Wiley (PUBLISHER • GB) |
| DOI | 10.1111/bjet.13494 |
| OpenAlex | W4399334961 |
| Language | EN |
| Citations received | 19 |
| References cited | 27 |
This study investigates the validity and reliability of generative large language models (LLMs), specifically ChatGPT and Google's Bard, in grading student essays in higher education based on an analytical grading rubric. A total of 15 experienced English as a foreign language (EFL) instructors and two LLMs were asked to evaluate three student essays of varying quality. The grading scale comprised five domains: grammar, content, organization, style & expression and mechanics. The results revealed that fine‐tuned ChatGPT model demonstrated a very high level of reliability with an intraclass correlation (ICC) score of 0.972, Default ChatGPT model exhibited an ICC score of 0.947 and Bard showed a substantial level of reliability with an ICC score of 0.919. Additionally, a significant overlap was observed in certain domains when comparing the grades assigned by LLMs and human raters. In conclusion, the findings suggest that while LLMs demonstrated a notable consistency and potential for grading competency, further fine‐tuning and adjustment are needed for a more nuanced understanding of non‐objective essay criteria. The study not only offers insights into the potential use of LLMs in grading student essays but also highlights the need for continued development and research. Practitioner notes What is already known about this topic Large language models (LLMs), such as OpenAI's ChatGPT and Google's Bard, are known for their ability to generate text that mimics human‐like conversation and writing. LLMs can perform various tasks, including essay grading. Intraclass correlation (ICC) is a statistical measure used to assess the reliability of ratings given by different raters (in this case, EFL instructors and LLMs). What this paper adds The study makes a unique contribution by directly comparing the grading performance of expert EFL instructors with two LLMs—ChatGPT and Bard—using an analytical grading scale. It provides robust empirical evidence showing high reliability of LLMs in grading essays, supported by high ICC scores. It specifically highlights that the overall efficacy of LLMs extends to certain domains of essay grading. Implications for practice and/or policy The findings open up potential new avenues for utilizing LLMs in academic settings, particularly for grading student essays, thereby possibly alleviating workload of educators. The paper's insistence on the need for further fine‐tuning of LLMs underlines the continual interplay between technological advancement and its practical applications. The results lay down a footprint for future research in advancing the use of AI in essay grading.
Conversation · Grading (engineering) · Grammar · Intraclass correlation · Linguistics · Mathematics education · Psychometrics · Rubric · Artificial Intelligence in Healthcare and Education · Clinical Psychology · Engineering · Psychology · Text Readability and Simplification · Topic Modeling
Exploring the dual impact of AI in post-entry language assessment
Factors Influencing University Students’ Intention to Use LLM Agents for Effective Self-Directed Learning
“Generative AI literacy across education and business
Can ChatGPT score ESL writing? A correlation analysis between teacher and GenAI scores
Source integration in source-based writing
Generative artificial intelligence for automated writing evaluation
Young L2 students’ use of an AI-assisted writing assessment and feedback tool
Evaluating GPT ratings of EFL writing
An investigation of utilizing technology in English as a foreign language writing
Generative Artificial Intelligence for Automated Qualitative Feedback
Automated analysis of common errors in L2 learner production
Roles of Generative Artificial Intelligence (GenAI) in English as a Foreign Language (EFL) Instruction
The promise and limits of artificial intelligence in writing assessment and instruction
Development and validation of a GPT-based rater for assessing communication skills using the Gap-Kalamazoo Communication Skills Assessment Form
From Technology‐Challenged Teachers to Empowered Digitalized Citizens
What Deserve Studying the Most? A Q‐Methodology Approach to Explore Stakeholders' Perspectives on Research Priorities in GenAI ‐Supported Second Language Education
Quality counts? Examining the role of feedback provider and feedback quality on students' feedback perceptions
Artificial Intelligence for Language Learning
Automated scoring in the era of artificial intelligence
Exploring the potential of using an AI language model for automated essay scoring
AI-generated feedback on writing
A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research
Explainable Automated Essay Scoring
More efficient processes for creating automated essay scoring frameworks
Writing evaluation
| Unique citing works | 19 |
|---|---|
| Citations per year | 19 |
| Citation span | 2025 - 2026 (2) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 18 |