Comparison of Three Methods to Generate Synthetic Datasets for Social Science
Li-jing Arthur Chang
Authors Information |
Citation |
Full Text |
Li-jing Arthur Chang
Department of Journalism and Media Studies, Jackson State University, Jackson, Mississippi, United States
Cite this paper as:Chang, L. A. (2025). Comparison of Three Methods to Generate Synthetic Datasets for Social Science.
Journal of Systemics, Cybernetics and Informatics, 23(7), 39-44. https://doi.org/10.54808/JSCI.23.07.39
Online ISSN (Journal): 1690-4524
Abstract
Many researchers often have difficulties finding enough data to test their hypotheses [1][2]. This study explores three different ways to create “synthetic”2 (i.e., artificial) data that mimics real-world data in statistical traits like correlations (i.e., relationships between the variables). To see how well these methods perform, the study compares the patterns of synthetic data to their real-world counterparts and sees how closely the data maintain the correlations. Additionally, the study uses seven machine learning3 prediction methods to see how these synthetic data perform. The findings indicate that two methods more effectively preserve the original correlation structure, while the third method yields better predictive performance.