Synthetic data as a method for increasing reproducibility and transparency in educational research
Bibliographic Data
| ID | 21714907 |
|---|---|
| Authors | Simon Grund (0000-0002-1290-8986), Oliver Lüdtke (0000-0001-9744-3059), Alexander Robitzsch (0000-0002-8226-3132) |
| Year | 2026 |
| Publication date | 2026-02-26 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Zeitschrift für Erziehungswissenschaft (JOURNAL) |
| Journal identifiers | ISSN: 1434-663X • E-ISSN: 1862-5215 |
| Publisher | Springer Science and Business Media LLC (PUBLISHER) |
| DOI | 10.1007/s11618-026-01396-6 |
| Language | EN |
| References cited | 64 |
Open data are often regarded as an important step towards improving the reproducibility and transparency of educational science. Yet, data sharing remains rare, and without open data, statistical analyses often remain irreproducible. In this article, we provide an introduction to synthetic data, a statistical technique based on multiple imputation (MI) that can be used to create simulated copies of the data that can be shared even when the original data cannot. To this end, we discuss reproducibility-related challenges of synthetic data and outline different approaches for generating synthetic data, including conventional and data-augmented MI (DA-MI) approaches to synthetic data. Furthermore, we conducted a case study using data from the PISA 2018 study, in which we aimed to address several challenges with synthetic data in educational research, such as missing data, multilevel data, and complex sampling designs. Our results indicate that these challenges can be addressed with relatively simple tools and that synthetic data can reproduce the results in a variety of statistical analyses. Finally, we discuss remaining challenges and directions for future research
Multiple Imputation of Missing Data for Multilevel Models
Integrative data analysis
Replicability, Robustness, and Reproducibility in Psychological Science
Lessons Learned from Pisa
Randomization-Based Inference about Latent Variables from Complex Samples
Umgang mit fehlenden Werten in der psychologischen Forschung
Promoting an open research culture
Missing data
Mice
Synthetic Datasets for Statistical Disclosure Control
Estimating the Prevalence of Transparency and Reproducibility-Related Research Practices in Psychology (2014–2017)
Multiple Imputation for Nonresponse in Surveys
Are We Wasting a Good Crisis? The Availability of Psychological Research Data after the Storm
The poor availability of psychological research data for reanalysis
How to protect privacy in open data
| Citation velocity | historical |
|---|---|
| Highly cited | No |