Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

A Survey on Evaluation of Large Language Models

Bibliographic Data

ID23322797
AuthorsYupeng Chang (0000-0001-7178-6088, Jilin University), Xu Wang (0000-0002-7634-6716, Jilin University), Jindong Wang (0000-0001-7040-9299, Microsoft Research Asia (China)), Yuan Wu (0000-0003-2167-2351, Jilin University), Linyi Yang (0000-0003-0667-7349, Westlake University), Kaijie Zhu (0009-0002-6220-1476, Chinese Academy of Sciences), Hao Chen (0000-0002-8873-8266, Carnegie Mellon University), Xiaoyuan Yi (0000-0003-2710-1613, Microsoft Research Asia (China)), Cunxiang Wang (0000-0002-3023-8082, Westlake University), Yidong Wang (0000-0001-6552-6466, Peking University), Wei Ye (0000-0001-7484-0609, Peking University), Yue Zhang (0000-0003-1198-882X, Westlake University), Yi Chang (0000-0003-3595-021X, Jilin University), Philip S Yu (0000-0002-3491-5968, University of Illinois Chicago), Qiang Yang (0000-0003-4210-9007, Hong Kong University of Science and Technology), Xing Xie (0000-0002-8608-8482, Microsoft Research Asia (China))
Year2024
Volume15
Issue3
Pages1-45
Publication date2024-06-30
Peer ReviewedYes
Open AccessYes
TypeARTICLE
VenueACM Transactions on Intelligent Systems and Technology (JOURNAL)
Journal identifiersISSN: 2157-6904 • E-ISSN: 2157-6912
PublisherAssociation for Computing Machinery (ACM) (PUBLISHER)
DOI10.1145/3641289
OpenAlexW4391136507
LanguageEN
Citations received102
References cited51

Large language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better understanding of their potential risks. Over the past years, significant efforts have been made to examine LLMs from various perspectives. This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate , where to evaluate , and how to evaluate . Firstly, we provide an overview from the perspective of evaluation tasks, encompassing general natural language processing tasks, reasoning, medical usage, ethics, education, natural and social sciences, agent applications, and other areas. Secondly, we answer the ‘where’ and ‘how’ questions by diving into the evaluation methods and benchmarks, which serve as crucial components in assessing the performance of LLMs. Then, we summarize the success and failure cases of LLMs in different tasks. Finally, we shed light on several future challenges that lie ahead in LLMs evaluation. Our aim is to offer invaluable insights to researchers in the realm of LLMs evaluation, thereby aiding the development of more proficient LLMs. Our key point is that evaluation should be treated as an essential discipline to better assist the development of LLMs. We consistently maintain the related open-source materials at: https://github.com/MLGroupJLU/LLM-eval-survey

Popularity · Artificial Intelligence in Healthcare and Education · Computer Science · Natural Language Processing Techniques · Psychology · Social Psychology · Topic Modeling

  • Can AI extort like humans do? Understanding the construction of illicit genres by large language models vs humans

    Open Access•Emily Chiang, Lily Calloway et al.•AI & Society•2026

  • Literature review on large language models (LLMs) for cross-cultural project management

    Open Access•Yutong Shen, Dilek Cetindamar et al.•AI & Society•2026

  • From delegation to moral abdication

    Open Access•Rainer Mühlhoff•AI & Society•2026

  • AI-Enabled governance for sustainable cities

    Open Access•Cao Hong, Chen-Hsiang Hong et al.•Sustainable Futures•2026

  • Unlocking Bias Detection

    Open Access•Shaina Raza, Oluwanifemi Bamgbose et al.•IEEE Transactions on Computational…•2024

  • Beyond Discrimination

    Open Access•Leda Tortora•Frontiers in Psychiatry•2024

  • From fragments to digital wholeness

    Open Access•L Cardarelli•Journal of Cultural Heritage•2024

  • Bias and Fairness in Large Language Models

    Open Access•Isabel O Gallegos, Ryan A Rossi et al.•Computational Linguistics•2024

  • Unveiling the multifaceted concept of cognitive security

    Open Access•Fran Casino•Technology in Society•2025

  • Epistemia in the classroom

    Open Access•Giulia Bini, Walter Quattrociocchi•AI & Society•2026

  • Predicting urban futures

    Open Access•Junxi Qu, Lingkun Meng et al.•Progress in Human Geography•2026

  • Acceptance of healthcare services based on the large language model in China

    Open Access•Haoze Li, Shimo Zhang et al.•BMC Public Health•2025

  • Governance of Generative AI

    Open Access•Araz Taeihagh•Policy and Society•2025

  • Dissociating language and thought in large language models

    Open Access•Kyle Mahowald, Anna A Ivanova et al.•Trends in Cognitive Sciences•2024

  • Natural Language Processing–Based Technologies Along the Customer Journey—A Systematic Review and Co‐Occurrence Analysis

    Open Access•Tom Ferber, Daryoush Daniel Vaziri et al.•Human Behavior and Emerging…•2025

  • Unlocking the multifaceted power of self-regulated learning and generative AI in foreign language morphological skills

    Open Access•Fangwei Huang, Haijing Zhang•Applied Linguistics•2026

  • GPT-4 generated psychological reports in psychodynamic perspective

    Open Access•Namwoo Kim, Jiseon Lee et al.•Frontiers in Psychiatry•2025

  • Visible AI assistant outputs in psychosocial risk management

    Open Access•Mehmed Zahid Çögenli•Frontiers in Public Health•2026

  • Generative Artificial Intelligence-driven orthodontic education practices

    Open Access•Menghan Zhang, Yuzhi Yang et al.•BMC Medical Education•2026

  • La incorporación de la inteligencia artificial generativa en la educación superior

    Open Access•Oriol Gómez-Soler, Juan González-Martínez•Edutec. Revista Electrónica de…•2026

  • A systematic review of human-LLM interactions in computational thinking empirical studies

    Yimei Zhang, Yajie Song et al.•Computer Science Education•2026

  • The Effects of AI Programming Assistant on University Students’ Algorithmic Thinking and Self-Efficacy

    Open Access•Wen Xiao, Xiao Wen et al.•Journal of Educational Computing…•2026

  • Leveraging large language models to assist philosophical counseling

    Open Access•Bokai Chen, Weiwei Zheng et al.•Humanities and Social Sciences…•2025

  • Large language models empowered agent-based modeling and simulation

    Open Access•Chen Gao, Xiaochong Lan et al.•Humanities and Social Sciences…•2024

  • Optimisation of university instructors’ professional practice through AI agents

    Open Access•Bayanali Doszhanov, Bakytzhan Zhumadilla et al.•Frontiers in Education•2026

  • LLM-generated competence-based e-assessment items for higher education mathematics

    Open Access•Roy Meissner, Alexander Pögelt et al.•Frontiers in Education•2024

  • Why are you traveling? Inferring trip profiles from online reviews and domain-knowledge

    Open Access•Lucas G S Félix, Washington Cunha et al.•Online Social Networks and Media•2025

  • Leveraging AI for Mental Healthcare in Social Fintech

    Open Access•Nguyen Khanh Son, Arun Kumar Sangaiah et al.•IEEE Transactions on Computational…•2025

  • Metaphors as Semantic Anchors

    Open Access•Hanqing Tao, Xuesong Wang et al.•IEEE Transactions on Computational…•2025

  • Glide

    Open Access•Yudan Lyu, Jingwei Lu et al.•IEEE Transactions on Computational…•2026

  • Analyzing Social Landscapes

    Open Access•Amirhossein Dezhboro, Pouria Babvey et al.•IEEE Transactions on Computational…•2026

  • Federated Service for Semantic Misalignment in Supply–Demand Matching

    Open Access•Shouwen Wang, Rui Qin et al.•IEEE Transactions on Computational…•2025

  • Advancing Team Conflict Detection by Leveraging Multifeature Embeddings, LMs, and LLMs

    Open Access•Astha Verma, Pratyush Yadav et al.•IEEE Transactions on Computational…•2026

  • Using Human Cumulative Prospect Theory to Understand Large Language Models Decision-Making

    Open Access•Guoshuai Zhang, Jiaji Wu et al.•IEEE Transactions on Computational…•2026

  • Evaluating TabPFN

    Open Access•Mazen Shawosh, Haseeb Nisar et al.•Frontiers in Public Health•2026

  • Exploring the Potential of Low-Barrier AI Tools for Culturally Responsive STEM Learning

    Open Access•Toiroa Williams, Minh Nguyen et al.•Education Sciences•2026

  • From Where to What

    Open Access•Richard Wen, Songnian Li•ISPRS International Journal of…•2026

  • Generative Artificial Intelligence in Geography

    Open Access•Sai Leung Ng, Chien-Min Chu•ISPRS International Journal of…•2026

  • On the Use of LLMs for GIS-Based Spatial Analysis

    Open Access•Roberto Pierdicca, Nikhil Muralikrishna et al.•ISPRS International Journal of…•2025

  • ChatGeoAI

    Open Access•Ali Mansourian, Rachid Oucheikh•ISPRS International Journal of…•2024

  • Government risk communication and response networks

    Open Access•Lihua Wang, Shengyi Jiang•International Journal of Disaster…•2025

  • Prioritise risks and improve adaptation strategies in the Veneto coast through the application of a custom AI tool

    Open Access•Maria Katherina Dal Barco, Veronica Casartelli et al.•International Journal of Disaster…•2025

  • Livestream communities for AI-generated video content

    Open Access•Sacha Gutierrez, Alkım Almila Akdağ Salah et al.•International Journal of…•2026

  • Enhancing user information disclosure intention in dynamic conversations of intelligent recommendation systems based on large language models

    Open Access•Chunze Xu, Fengqiang Gao et al.•International Journal of…•2025

  • Application of multi-agent systems in legal education

    Open Access•S J Shi, Ying Cao et al.•Interactive Learning Environments•2026

  • Exploring the depression sharing on social media platforms

    Open Access•Keke Hou, Tingting Hou et al.•Current Psychology•2025

  • LLM-aided representation of human commuting patterns in 42 Chinese cities with open data

    Open Access•Enyuan Cao, Justin Hayse Chiwing G Tang et al.•Journal of Transport Geography•2026

  • Partisan Knowledge Claims in Congressional Oversight

    Open Access•Kenneth Lowande, Mark A Weiss•Legislative Studies Quarterly•2026

  • Investigating the effects of an LLM-based Socratic conversational agent on students’ academic performance and reflective thinking in higher education

    Open Access•Linjin Xi, Yi Zhang et al.•Computers & Education•2026

  • Exploring the prospects of multimodal large language models for Automated Emotion Recognition in education

    Open Access•Shuzhen Yu, Alexey Androsov et al.•Computers & Education•2025

  • Boosting Student Engagement in STEM

    Open Access•Minkai Wang, Jingdong Zhu et al.•Journal of Computer Assisted…•2025

  • Research note

    Open Access•Yu-Hsin Tung, Zhe-Rui Yang et al.•Landscape and Urban Planning•2025

  • Enhancing traditional surveys with village view imagery and multimodal large language models

    Open Access•Muzhe Pan, Weipan Xu et al.•Applied Geography•2026

  • Unpacking Beliefs and Engagement in AI ‐Assisted Chinese Learning

    Open Access•Xuesong Zhai, Chen Wu et al.•European Journal of Education•2025

  • Opinion paper

    Open Access•Benedetto Lepori, Jens Peter Andersen et al.•Scientometrics•2026

  • Detecting AI adoption at scale

    Open Access•Ana Pastor-Merino, Xavier Martínez-Barbero et al.•Scientometrics•2026

  • Large language models in peer review

    Open Access•Zhuanlan Sun•Scientometrics•2025

  • Who gets the money? A qualitative analysis of fintech lending and credit scoring through the adoption of AI and alternative data

    Open Access•Maximilian Tigges, Sönke Mestwerdt et al.•Technological Forecasting and…•2024

  • A hybrid mixed methods design of qualitative enhancement and reciprocal feedback loop for augmented text classification

    Open Access•Gahl Silverman, Dov Te’eni et al.•Quality & Quantity•2025

  • Assessing student perceptions and use of instructor versus AI ‐generated feedback

    Open Access•Erkan Er, Gökhan Akçapınar et al.•British Journal of Educational…•2025

  • The interplay of learning, analytics and artificial intelligence in education

    Open Access•Mutlu Cukurova•British Journal of Educational…•2025

  • Cognitive Echo

    Open Access•Longwei Zheng, Anna He et al.•British Journal of Educational…•2025

  • Leveraging LLM respondents for item evaluation

    Open Access•Yunting Liu, Shreya Bhandari et al.•British Journal of Educational…•2025

  • Ink and algorithm

    Open Access•Kaixun Yang, Yixin Cheng et al.•British Journal of Educational…•2026

  • Comparative analysis of GPT-4, Gemini, and Ernie as gloss sign language translators in special education

    Open Access•Achraf Othman, Khansa Chemnad et al.•Discover Global Society•2024

  • Geo-hallucination in urban analytics

    Open Access•Xiao Huang•Environment and Planning B Urban…•2025

  • Exploring the potential of large language models (LLMs) in analyzing passengers’ perceptions of transit service quality

    Open Access•Shuli Luo, Sylvia Y He et al.•Environment and Planning B Urban…•2026

  • Empowering SMEs with SustainWater Bot to advance urban water sustainability

    Open Access•Muhammad Arslan, Saba Munawar et al.•Sustainable Cities and Society•2025

  • Deep learning for optimizing urban governance by "sensing-processing-responding" cycle

    Open Access•Mingjun Cheng, Hong Jin et al.•Sustainable Cities and Society•2025

  • A novel domain knowledge augmented large language model based medical conversation system for sustainable smart city development

    Open Access•Haochen Zou, Yongli Wang et al.•Sustainable Cities and Society•2025

  • Generative AI in academic writing

    Open Access•Sameh Kamal Mohamed Ibrahim, Zakaria Abdelaziz Zakaria Mahmoud•Humanities and Social Sciences…•2026

  • Between Innovation and Tradition

    Open Access•Alma Espartinez•Social Sciences•2025

  • Generative Artificial Intelligence and Regulations

    Open Access•Matteo Bodini•Societies•2024

  • Attributing Mind to Large Language Models

    Open Access•Oliver Jacobs, Oliver L Jacobs et al.•International Journal of Social…•2026

  • Between regulation and accessibility

    Open Access•Qin Xie, Ming Li et al.•Globalisation Societies and…•2025

  • Political Debate

    Open Access•Michael Burnham, Kayla Kahn et al.•Political Analysis•2025

  • Framing of higher education opportunities by large language models

    Open Access•Pii-Tuulia Nikula•AI & Society•2026

  • Minds of their own? Decoding free will in large language models

    Open Access•Andrea Lavazza, Giuseppe Sartori et al.•AI & Society•2026

  • LLM-based NLG Evaluation

    Open Access•Mingqi Gao, Xinyu Hu et al.•Computational Linguistics•2025

  • Compositionality and Sentence Meaning

    Open Access•James Fodor, Simon De Deyne et al.•Computational Linguistics•2024

  • Automating Plan Evaluation Using Agentic Large Language Models

    Open Access•Xinyu Fu, Chaosu Li•Journal of Planning Education and…•2025

  • GISedu-GPT

    Zhiyun Wang, Yifan Zhang et al.•Journal of Geography in Higher…•2025

  • Generative AI as a Knowledge Distribution System

    Andrea Lavazza, Mirko Farina•Social Epistemology•2026

  • When AI sees hotter

    Open Access•Tenzin Tamang, Ruilin Zheng•Public Understanding of Science•2026

  • A survey on moral foundation theory and pre-trained language models

    Open Access•Lorenzo Zangari, Candida M Greco et al.•AI & Society•2025

  • Using Artificial Intelligence to Support Scientometric Analysis of Scholarly Literature

    Open Access•H Luan, Brian E Perron et al.•Journal of the Society for Social…•2024

  • Act as an expert in psychometry. The evaluation of large language models utility in psychological tests cross-cultural adaptations

    Open Access•Jarosław Grobelny, Kacper Szymański et al.•Acta Psychologica•2025

  • The bewitching AI

    Open Access•Emanuele Bottazzi Grifoni, Roberta Ferrario•Philosophy & Technology•2025

  • ChatGPT as a Patient Education Tool in Female Sexual Dysfunction

    Open Access•Celal Akdemir, Mücahit Furkan Balcı et al.•International Journal of Sexual…•2026

  • Cognitive Automation and Sustainable Development

    Open Access•Yuanfan Li, Rongrong Li et al.•Sustainable Development•2026

  • Large language confusion

    Open Access•Andrey S Druzhinin, Diego A Ramírez•Language Sciences•2026

  • Natural language processing for social science research

    Open Access•Yuxin Hou, Junming Huang•Chinese Journal of Sociology•2025

  • A large language model-based approach to building the knowledge graph for master plan

    Open Access•W S Zhang, Wusiqin Zhang et al.•Land Use Policy•2025

  • Leveraging large language models for tourism research based on 5D framework

    Open Access•Jin Rui, Yuhan Xu et al.•Tourism Management•2025

  • Administrative Decision-Making with Generative AI

    Open Access•Yushim Kim, Jieun Kim et al.•Administration & Society•2026

  • Stress detection through prompt engineering with a general-purpose LLM

    Open Access•Nima Esmi, Asadollah Shahbahrami et al.•Acta Psychologica•2025

  • The Epistemic Impact of Large Language Models on Policymaking

    Open Access•Naikang Feng, Naishi Feng et al.•Policy Studies Journal•2026

  • Does AI boost firm productivity? A web scraping and LLMs approach

    Open Access•Ana Pastor-Merino, Xavier Martínez-Barbero et al.•Telecommunications Policy•2026

  • XunZi-MLLM

    Open Access•Dongmei Zhu, Chang Liu et al.•Digital Scholarship in the…•2025

  • Intelligent Computing Social Modeling and Methodological Innovations in Political Science in the Era of Large Language Models

    Open Access•Zhenyu Wang, Dequan Wang et al.•Journal of Chinese Political…•2025

  • Advances In Experimental Social Psychology

    Bertram Gawronski•Advances in Experimental Social…•2022

  • Evaluating the Feasibility of ChatGPT in Healthcare

    Open Access•Marco Cascella, Jonathan Montomoli et al.•Journal of Medical Systems•2023

  • GPT-3

    Open Access•Luciano Floridi, Massimo Chiriatti•Minds and Machines•2020

  • ImageNet

    Jia Deng, Wei Dong et al.•2009 IEEE Conference on Computer…•2009

  • How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment

    Open Access•Aidan Gilson, Conrad W Safranek et al.•JMIR Medical Education•2023

  • Support-Vector Networks

    Open Access•Corinna Cortes, Vladimir Vapnik•Machine Learning•1995

  • Bart

    Open Access•Mike Lewis, Yinhan Liu et al.•Proceedings of the 58th Annual…•2020

  • Validity problems comparing values across cultures and possible solutions.

    Kaiping Peng, Richard E Nisbett et al.•Psychological Methods•1997

  • Microsoft Coco

    Open Access•Tsung-Yi Lin, Michael Maire et al.•Computer Vision -- ECCV 2014•2014

  • Performance of ChatGPT on USMLE

    Open Access•Tiffany H Kung, Morgan Cheatham et al.•PLOS Digital Health•2023

  • Deep learning

    Open Access•Yann LeCun, Yoshua Bengio et al.•Nature•2015

  • ChatGPT for good? On opportunities and challenges of large language models for education

    Open Access•Enkelejda Kasneci, Kathrin Sessler et al.•Learning and Individual Differences•2023

  • The Self-Perception and Political Biases of ChatGPT

    Open Access•Jérôme Rutinowski, Sven Franke et al.•Human Behavior and Emerging…•2024

  • Online question and answer sessions

    Open Access•Malin Jansson, Stefan Hrastinski et al.•The Internet and Higher Education•2021

  • Emotional intelligence of Large Language Models

    Open Access•Xuena Wang, Xueting Li et al.•Journal of Pacific Rim Psychology•2023

Unique citing works102
Citations per year51
Citation span2024 - 2026 (3)
Citation velocitycurrent
Highly citedYes
Citation typesNeutral: 91

Tools

Open DOI
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae