Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

Utilizing large language models for EFL essay grading

An examination of reliability and validity in rubric‐based assessments

Bibliographic Data

ID21297215
AuthorsFatih Yavuz (0000-0003-2645-2710, Preparatory Department Mudanya University Mudanya Turkey, corresponding author), Özgür Çelik (0000-0002-0300-9073, School of Foreign Languages Balıkesir University Balıkesir Turkey), Gamze Yavaş Çelik (0000-0003-1571-9686, School of Foreign Languages Balıkesir University Balıkesir Turkey)
Year2025
Volume56
Issue1
Pages150-166
Publication date2025-01-01
Peer ReviewedYes
Open AccessYes
TypeARTICLE
VenueBritish Journal of Educational Technology (JOURNAL)
Journal identifiersISSN: 0007-1013 • E-ISSN: 1467-8535
PublisherWiley (PUBLISHER • GB)
DOI10.1111/bjet.13494
OpenAlexW4399334961
LanguageEN
Citations received19
References cited27

This study investigates the validity and reliability of generative large language models (LLMs), specifically ChatGPT and Google's Bard, in grading student essays in higher education based on an analytical grading rubric. A total of 15 experienced English as a foreign language (EFL) instructors and two LLMs were asked to evaluate three student essays of varying quality. The grading scale comprised five domains: grammar, content, organization, style & expression and mechanics. The results revealed that fine‐tuned ChatGPT model demonstrated a very high level of reliability with an intraclass correlation (ICC) score of 0.972, Default ChatGPT model exhibited an ICC score of 0.947 and Bard showed a substantial level of reliability with an ICC score of 0.919. Additionally, a significant overlap was observed in certain domains when comparing the grades assigned by LLMs and human raters. In conclusion, the findings suggest that while LLMs demonstrated a notable consistency and potential for grading competency, further fine‐tuning and adjustment are needed for a more nuanced understanding of non‐objective essay criteria. The study not only offers insights into the potential use of LLMs in grading student essays but also highlights the need for continued development and research. Practitioner notes What is already known about this topic Large language models (LLMs), such as OpenAI's ChatGPT and Google's Bard, are known for their ability to generate text that mimics human‐like conversation and writing. LLMs can perform various tasks, including essay grading. Intraclass correlation (ICC) is a statistical measure used to assess the reliability of ratings given by different raters (in this case, EFL instructors and LLMs). What this paper adds The study makes a unique contribution by directly comparing the grading performance of expert EFL instructors with two LLMs—ChatGPT and Bard—using an analytical grading scale. It provides robust empirical evidence showing high reliability of LLMs in grading essays, supported by high ICC scores. It specifically highlights that the overall efficacy of LLMs extends to certain domains of essay grading. Implications for practice and/or policy The findings open up potential new avenues for utilizing LLMs in academic settings, particularly for grading student essays, thereby possibly alleviating workload of educators. The paper's insistence on the need for further fine‐tuning of LLMs underlines the continual interplay between technological advancement and its practical applications. The results lay down a footprint for future research in advancing the use of AI in essay grading.

Conversation · Grading (engineering) · Grammar · Intraclass correlation · Linguistics · Mathematics education · Psychometrics · Rubric · Artificial Intelligence in Healthcare and Education · Clinical Psychology · Engineering · Psychology · Text Readability and Simplification · Topic Modeling

  • Exploring the dual impact of AI in post-entry language assessment

    Tiancheng Zhang, Rosemary Erlam et al.•Annual Review of Applied…•2025

  • Factors Influencing University Students’ Intention to Use LLM Agents for Effective Self-Directed Learning

    Open Access•Jiarong Fan, Jiahao Fan et al.•TechTrends•2026

  • “Generative AI literacy across education and business

    Open Access•Michael Reicho, Kathrin Otrel-Cass et al.•International Journal of…•2026

  • Can ChatGPT score ESL writing? A correlation analysis between teacher and GenAI scores

    Open Access•Anas Alkhofi•Language Teaching Research•2026

  • Source integration in source-based writing

    Open Access•Andrew Potter, Yu Tian et al.•Assessing Writing•2026

  • Generative artificial intelligence for automated writing evaluation

    Open Access•Shadi I Abudalfa, Jessie S Barrot•Assessing Writing•2026

  • Young L2 students’ use of an AI-assisted writing assessment and feedback tool

    Open Access•Mikyung Kim Wolf, Michael Suhan et al.•Assessing Writing•2026

  • Evaluating GPT ratings of EFL writing

    Open Access•Yi Chen•Assessing Writing•2026

  • An investigation of utilizing technology in English as a foreign language writing

    Open Access•Xincheng Wu•Reading and Writing•2026

  • Generative Artificial Intelligence for Automated Qualitative Feedback

    Open Access•Jessie S Barrot, Hung Phu Bui•RELC Journal•2026

  • Automated analysis of common errors in L2 learner production

    Open Access•Atsushi Mizumoto•Studies in Second Language…•2025

  • Roles of Generative Artificial Intelligence (GenAI) in English as a Foreign Language (EFL) Instruction

    Open Access•Luying Deng, Khairul Azhar Jamaludin•SAGE Open•2026

  • The promise and limits of artificial intelligence in writing assessment and instruction

    Open Access•Yue-liang Pan, Yue‐Liang Pan et al.•Educational Technology Research…•2026

  • Development and validation of a GPT-based rater for assessing communication skills using the Gap-Kalamazoo Communication Skills Assessment Form

    Yu-Jeng Ju, Yi-Ching Wang et al.•Medical Teacher•2026

  • From Technology‐Challenged Teachers to Empowered Digitalized Citizens

    Open Access•Ziwen Pan, Yongliang Wang•European Journal of Education•2025

  • What Deserve Studying the Most? A Q‐Methodology Approach to Explore Stakeholders' Perspectives on Research Priorities in GenAI ‐Supported Second Language Education

    Open Access•Ran Zhi, Ziwen Pan•European Journal of Education•2025

  • Quality counts? Examining the role of feedback provider and feedback quality on students' feedback perceptions

    Open Access•Theresa Ruwe, Livia Kuklick•British Journal of Educational…•2026

  • Artificial Intelligence for Language Learning

    Open Access•Shen Qiao, Mingyue Gu et al.•International Journal of Applied…•2025

  • Automated scoring in the era of artificial intelligence

    Open Access•Burak Aydın, Tarık Kışla et al.•System•2025

  • Exploring the potential of using an AI language model for automated essay scoring

    Open Access•Atsushi Mizumoto, Masaki Eguchi•Research Methods in Applied…•2023

  • AI-generated feedback on writing

    Open Access•Juan Escalante, Austin Pack et al.•International Journal of…•2023

  • A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research

    Open Access•Terry K Koo, Mae Y Li•Journal of Chiropractic Medicine•2016

  • Explainable Automated Essay Scoring

    Open Access•Vive Kumar, Vivekanandan Kumar et al.•Frontiers in Education•2020

  • More efficient processes for creating automated essay scoring frameworks

    Open Access•Jinnie Shin, Mark J Gierl•Language Testing•2021

  • Writing evaluation

    Open Access•Nahla Nola Bacha, Nahla Bacha•System•2001

Unique citing works19
Citations per year19
Citation span2025 - 2026 (2)
Citation velocitycurrent
Highly citedNo
Citation typesNeutral: 18

Tools

Open DOIOpen Access
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae