Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

Evaluating human versus machine learning performance in classifying research abstracts

Bibliographic Data

ID21443515
AuthorsYeow Chong Goh (Nanyang Technological University), Xin Cai (0000-0002-7191-2571, Nanyang Technological University), Xin Qing Cai, Walter Theseira (0000-0002-8738-2341, Singapore University of Social Sciences), Giovanni Ko (Singapore Management University), Khiam Aik Khor (0000-0003-1954-8423, Nanyang Technological University, corresponding author)
Year2020
Volume125
Issue2
Pages1197-1212
Publication date2020-11-01
Peer ReviewedYes
Open AccessYes
TypeARTICLE
VenueScientometrics (JOURNAL)
Journal identifiersISSN: 0138-9130 • E-ISSN: 1588-2861
PublisherSpringer Science and Business Media LLC (PUBLISHER)
DOI10.1007/s11192-020-03614-2
PMID32836529
OpenAlexW3043669475
LanguageEN
Citations received9
References cited31

We study whether humans or machine learning (ML) classification models are better at classifying scientific research abstracts according to a fixed set of discipline groups. We recruit both undergraduate and postgraduate assistants for this task in separate stages, and compare their performance against the support vectors machine ML algorithm at classifying European Research Council Starting Grant project abstracts to their actual evaluation panels, which are organised by discipline groups. On average, ML is more accurate than human classifiers, across a variety of training and test datasets, and across evaluation panels. ML classifiers trained on different training sets are also more reliable than human classifiers, meaning that different ML classifiers are more consistent in assigning the same classifications to any given abstract, compared to different human classifiers. While the top five percentile of human classifiers can outperform ML in limited cases, selection and training of such classifiers is likely costly and difficult compared to training ML models. Our results suggest ML models are a cost effective and highly accurate method for addressing problems in comparative bibliometric analysis, such as harmonising the discipline classifications of research from different funding agencies or countries

Machine learning · Meaning (existential) · Percentile · Set (abstract data type) · Statistics · Task (project management) · Training set · Variety (cybernetics) · Advanced Text Analysis Techniques · Artificial Intelligence · Biomedical Text Mining and Ontologies · Computer Science · Engineering · Mathematics · Psychology · Topic Modeling

  • Using Natural Language Processing to Identify Stigmatizing Language in Labor and Birth Clinical Notes

    Open Access•Veronica Barcelona, Danielle Scharp et al.•Maternal and Child Health Journal•2023

  • Essential signals in publication trends and collaboration patterns in global Research Integrity and Research Ethics (Rire)

    Open Access•A M Soehartono, L G Yu et al.•Scientometrics•2022

  • Journal article classification using abstracts

    Open Access•Cristina Arhiliuc, Raf Guns et al.•Scientometrics•2025

  • Automatic noise reduction of domain-specific bibliographic datasets using positive-unlabeled learning

    Open Access•Guo Chen, Jing Chen et al.•Scientometrics•2023

  • A transfer learning approach to interdisciplinary document classification with keyword-based explanation

    Open Access•Xiaoming Huang, Peihu Zhu et al.•Scientometrics•2023

  • When something goes wrong

    Open Access•Andrea Berber, Sanja Srećković•AI & Society•2024

  • Critically Engaged Pragmatism

    Open Access•Carole J Lee•Social Epistemology•2026

  • Ten Essential Pillars in Artificial Intelligence for University Science Education

    Open Access•Angel Deroncele-Acosta, Omar Bellido-Valdiviezo et al.•SAGE Open•2024

  • Supercomputers and quantum computing on the axis of cyber security

    Open Access•Haydar Yalçın, Haydar Yalcin et al.•Technology in Society•2024

  • An Application of Hierarchical Kappa-type Statistics in the Assessment of Majority Agreement among Multiple Observers

    J R Landis, Gary G Koch•Biometrics•1977

  • Support-Vector Networks

    Open Access•Corinna Cortes, Vladimir Vapnik•Machine Learning•1995

  • Universality of citation distributions

    Open Access•Filippo Radicchi, Santo Fortunato et al.•Proceedings of the National…•2008

  • A systematic analysis of performance measures for classification tasks

    Open Access•Marina Sokolova, Guy Lapalme•Information Processing & Management•2009

  • The scientific impact of nations

    Open Access•David A King•Nature•2004

  • Co‐citation in the scientific literature

    Open Access•Henry Small•Journal of the American Society…•1973

  • Interrater reliability

    Open Access•Marry L McHugh•Biochemia Medica•2012

  • Measuring nominal scale agreement among many raters

    Joseph L Flei, Joseph L Fleiss•Psychological Bulletin•1971

  • From translations to problematic networks

    Open Access•Michel Callon, Jean-Pierre Courtial et al.•Social Science Information•1983

Unique citing works9
Citations per year2,25
Citation span2022 - 2026 (5)
Citation velocitycurrent
Highly citedNo
Citation typesNeutral: 9

Tools

Open DOIOpen Access
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae