A Survey on Evaluation of Large Language Models
Bibliographic Data
| ID | 23322797 |
|---|---|
| Authors | Yupeng Chang (0000-0001-7178-6088, Jilin University), Xu Wang (0000-0002-7634-6716, Jilin University), Jindong Wang (0000-0001-7040-9299, Microsoft Research Asia (China)), Yuan Wu (0000-0003-2167-2351, Jilin University), Linyi Yang (0000-0003-0667-7349, Westlake University), Kaijie Zhu (0009-0002-6220-1476, Chinese Academy of Sciences), Hao Chen (0000-0002-8873-8266, Carnegie Mellon University), Xiaoyuan Yi (0000-0003-2710-1613, Microsoft Research Asia (China)), Cunxiang Wang (0000-0002-3023-8082, Westlake University), Yidong Wang (0000-0001-6552-6466, Peking University), Wei Ye (0000-0001-7484-0609, Peking University), Yue Zhang (0000-0003-1198-882X, Westlake University), Yi Chang (0000-0003-3595-021X, Jilin University), Philip S Yu (0000-0002-3491-5968, University of Illinois Chicago), Qiang Yang (0000-0003-4210-9007, Hong Kong University of Science and Technology), Xing Xie (0000-0002-8608-8482, Microsoft Research Asia (China)) |
| Year | 2024 |
| Volume | 15 |
| Issue | 3 |
| Pages | 1-45 |
| Publication date | 2024-06-30 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | ACM Transactions on Intelligent Systems and Technology (JOURNAL) |
| Journal identifiers | ISSN: 2157-6904 • E-ISSN: 2157-6912 |
| Publisher | Association for Computing Machinery (ACM) (PUBLISHER) |
| DOI | 10.1145/3641289 |
| OpenAlex | W4391136507 |
| Language | EN |
| Citations received | 102 |
| References cited | 51 |
Large language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better understanding of their potential risks. Over the past years, significant efforts have been made to examine LLMs from various perspectives. This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate , where to evaluate , and how to evaluate . Firstly, we provide an overview from the perspective of evaluation tasks, encompassing general natural language processing tasks, reasoning, medical usage, ethics, education, natural and social sciences, agent applications, and other areas. Secondly, we answer the ‘where’ and ‘how’ questions by diving into the evaluation methods and benchmarks, which serve as crucial components in assessing the performance of LLMs. Then, we summarize the success and failure cases of LLMs in different tasks. Finally, we shed light on several future challenges that lie ahead in LLMs evaluation. Our aim is to offer invaluable insights to researchers in the realm of LLMs evaluation, thereby aiding the development of more proficient LLMs. Our key point is that evaluation should be treated as an essential discipline to better assist the development of LLMs. We consistently maintain the related open-source materials at: https://github.com/MLGroupJLU/LLM-eval-survey
Popularity · Artificial Intelligence in Healthcare and Education · Computer Science · Natural Language Processing Techniques · Psychology · Social Psychology · Topic Modeling
Can AI extort like humans do? Understanding the construction of illicit genres by large language models vs humans
Literature review on large language models (LLMs) for cross-cultural project management
From delegation to moral abdication
AI-Enabled governance for sustainable cities
Unlocking Bias Detection
Beyond Discrimination
From fragments to digital wholeness
Bias and Fairness in Large Language Models
Unveiling the multifaceted concept of cognitive security
Epistemia in the classroom
Predicting urban futures
Acceptance of healthcare services based on the large language model in China
Governance of Generative AI
Dissociating language and thought in large language models
Natural Language Processing–Based Technologies Along the Customer Journey—A Systematic Review and Co‐Occurrence Analysis
Unlocking the multifaceted power of self-regulated learning and generative AI in foreign language morphological skills
GPT-4 generated psychological reports in psychodynamic perspective
Visible AI assistant outputs in psychosocial risk management
Generative Artificial Intelligence-driven orthodontic education practices
La incorporación de la inteligencia artificial generativa en la educación superior
A systematic review of human-LLM interactions in computational thinking empirical studies
The Effects of AI Programming Assistant on University Students’ Algorithmic Thinking and Self-Efficacy
Leveraging large language models to assist philosophical counseling
Large language models empowered agent-based modeling and simulation
Optimisation of university instructors’ professional practice through AI agents
LLM-generated competence-based e-assessment items for higher education mathematics
Why are you traveling? Inferring trip profiles from online reviews and domain-knowledge
Leveraging AI for Mental Healthcare in Social Fintech
Metaphors as Semantic Anchors
Glide
Analyzing Social Landscapes
Federated Service for Semantic Misalignment in Supply–Demand Matching
Advancing Team Conflict Detection by Leveraging Multifeature Embeddings, LMs, and LLMs
Using Human Cumulative Prospect Theory to Understand Large Language Models Decision-Making
Evaluating TabPFN
Exploring the Potential of Low-Barrier AI Tools for Culturally Responsive STEM Learning
From Where to What
Generative Artificial Intelligence in Geography
On the Use of LLMs for GIS-Based Spatial Analysis
ChatGeoAI
Government risk communication and response networks
Prioritise risks and improve adaptation strategies in the Veneto coast through the application of a custom AI tool
Livestream communities for AI-generated video content
Enhancing user information disclosure intention in dynamic conversations of intelligent recommendation systems based on large language models
Application of multi-agent systems in legal education
Exploring the depression sharing on social media platforms
LLM-aided representation of human commuting patterns in 42 Chinese cities with open data
Partisan Knowledge Claims in Congressional Oversight
Investigating the effects of an LLM-based Socratic conversational agent on students’ academic performance and reflective thinking in higher education
Exploring the prospects of multimodal large language models for Automated Emotion Recognition in education
Boosting Student Engagement in STEM
Research note
Enhancing traditional surveys with village view imagery and multimodal large language models
Unpacking Beliefs and Engagement in AI ‐Assisted Chinese Learning
Opinion paper
Detecting AI adoption at scale
Large language models in peer review
Who gets the money? A qualitative analysis of fintech lending and credit scoring through the adoption of AI and alternative data
A hybrid mixed methods design of qualitative enhancement and reciprocal feedback loop for augmented text classification
Assessing student perceptions and use of instructor versus AI ‐generated feedback
The interplay of learning, analytics and artificial intelligence in education
Cognitive Echo
Leveraging LLM respondents for item evaluation
Ink and algorithm
Comparative analysis of GPT-4, Gemini, and Ernie as gloss sign language translators in special education
Geo-hallucination in urban analytics
Exploring the potential of large language models (LLMs) in analyzing passengers’ perceptions of transit service quality
Empowering SMEs with SustainWater Bot to advance urban water sustainability
Deep learning for optimizing urban governance by "sensing-processing-responding" cycle
A novel domain knowledge augmented large language model based medical conversation system for sustainable smart city development
Generative AI in academic writing
Between Innovation and Tradition
Generative Artificial Intelligence and Regulations
Attributing Mind to Large Language Models
Between regulation and accessibility
Political Debate
Framing of higher education opportunities by large language models
Minds of their own? Decoding free will in large language models
LLM-based NLG Evaluation
Compositionality and Sentence Meaning
Automating Plan Evaluation Using Agentic Large Language Models
GISedu-GPT
Generative AI as a Knowledge Distribution System
When AI sees hotter
A survey on moral foundation theory and pre-trained language models
Using Artificial Intelligence to Support Scientometric Analysis of Scholarly Literature
Act as an expert in psychometry. The evaluation of large language models utility in psychological tests cross-cultural adaptations
The bewitching AI
ChatGPT as a Patient Education Tool in Female Sexual Dysfunction
Cognitive Automation and Sustainable Development
Large language confusion
Natural language processing for social science research
A large language model-based approach to building the knowledge graph for master plan
Leveraging large language models for tourism research based on 5D framework
Administrative Decision-Making with Generative AI
Stress detection through prompt engineering with a general-purpose LLM
The Epistemic Impact of Large Language Models on Policymaking
Does AI boost firm productivity? A web scraping and LLMs approach
XunZi-MLLM
Intelligent Computing Social Modeling and Methodological Innovations in Political Science in the Era of Large Language Models
Advances In Experimental Social Psychology
Evaluating the Feasibility of ChatGPT in Healthcare
GPT-3
ImageNet
How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment
Support-Vector Networks
Bart
Validity problems comparing values across cultures and possible solutions.
Microsoft Coco
Performance of ChatGPT on USMLE
Deep learning
ChatGPT for good? On opportunities and challenges of large language models for education
The Self-Perception and Political Biases of ChatGPT
Online question and answer sessions
Emotional intelligence of Large Language Models
| Unique citing works | 102 |
|---|---|
| Citations per year | 51 |
| Citation span | 2024 - 2026 (3) |
| Citation velocity | current |
| Highly cited | Yes |
| Citation types | Neutral: 91 |