Assessing Replicability of Machine Learning Results
An Introduction to Methods on Predictive Accuracy in Social Sciences
Bibliographic Data
| ID | 12171069 |
|---|---|
| Authors | Ranjith Vijayakumar (National University of Singapore, Singapore), Mike W-L Cheung (0000-0003-0113-0758, National University of Singapore, Singapore) |
| Year | 2019 |
| Volume | 39 |
| Issue | 5 |
| Pages | 768-801 |
| Publication date | 2019-12-09 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Social Science Computer Review (JOURNAL) |
| Journal identifiers | ISSN: 0894-4393 • E-ISSN: 1552-8286 |
| Publisher | SAGE Publishing (PUBLISHER • US) |
| DOI | 10.1177/0894439319888445 |
| OpenAlex | W2995559771 |
| Language | EN |
| Citations received | 1 |
| References cited | 51 |
Machine learning methods have become very popular in diverse fields due to their focus on predictive accuracy, but little work has been conducted on how to assess the replicability of their findings. We introduce and adapt replication methods advocated in psychology to the aims and procedural needs of machine learning research. In Study 1, we illustrate these methods with the use of an empirical data set, assessing the replication success of a predictive accuracy measure, namely, R 2 on the cross-validated and test sets of the samples. We introduce three replication aims. First, tests of inconsistency examine whether single replications have successfully rejected the original study. Rejection will be supported if the 95% confidence interval (CI) of R 2 difference estimates between replication and original does not contain zero. Second, tests of consistency help support claims of successful replication. We can decide apriori on a region of equivalence, where population values of the difference estimates are considered equivalent for substantive reasons. The 90% CI of a different estimate lying fully within this region supports replication. Third, we show how to combine replications to construct meta-analytic intervals for better precision of predictive accuracy measures. In Study 2, R 2 is reduced from the original in a subset of replication studies to examine the ability of the replication procedures to distinguish true replications from nonreplications. We find that when combining studies sampled from same population to form meta-analytic intervals, random-effects methods perform best for cross-validated measures while fixed-effects methods work best for test measures. Among machine learning methods, regression was comparable to many complex methods, while support vector machine performed most reliably across a variety of scenarios. Social scientists who use machine learning to model empirical data can use these methods to enhance the reliability of their findings
Confidence interval · Consistency (knowledge bases · Empirical research · Equivalence (formal languages · Machine learning · Meta-analysis · Population · Replication (statistics · Set (abstract data type · Statistics · Computational and Text Analysis Methods · Computer Science · Data Analysis with R · Mathematics · Medicine · Mental Health Research Topics · Artificial Intelligence
Meta‐Analysis
Bootstrap Methods and their Application
Modern Applied Statistics with S
Introduction to Meta‐Analysis
Choosing Prediction Over Explanation in Psychology
Fixed- and random-effects models in meta-analysis.
The Elements of Statistical Learning
The Statistical Crisis in Science
To Explain or to Predict?
Toward using confidence intervals to compare correlations.
Building Predictive Models in R Using the caret Package
Regularization Paths for Generalized Linear Models via Coordinate Descent
False-Positive Psychology
Equivalence Tests
Bootstrap Standard Error and Confidence Intervals for the Difference Between Two Squared Multiple Correlation Coefficients
Predictive Accuracy as an Achievable Goal of Science
Are There Universal Aspects in the Structure and Contents of Human Values
Inference by Eye
Is psychology suffering from a replication crisis? What does “failure to replicate” really mean
An Overview of the Schwartz Theory of Basic Values
Big Data
Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015
Correlations redux
| Unique citing works | 1 |
|---|---|
| Citations per year | 0,25 |
| Citation span | 2022 - 2022 (1) |
| Citation velocity | historical |
| Highly cited | No |
| Citation types | Neutral: 1 |