Conditional Tabular Generative Adversarial Network-based Synthetic Data Generation for Model Generalisation Improvement
DOI:
https://doi.org/10.32890/jict2026.25.1.1Keywords:
Deep learning, CTGAN, data augmentation, synthetic dataAbstract
Accessing extensive and varied datasets is essential for developing strong predictive models in data analytics. However, many real-world applications suffer from small and imbalanced datasets, leading to overfitting, poor generalisation, and low model performance. Traditional data augmentation techniques are often unsuitable for tabular data, as they fail to preserve complex feature relationships. To address this challenge, this study adapts the Conditional Tabular Generative Adversarial Network (CTGAN) for synthetic data generation. The proposed approach involves five phases: (1) Data Acquisition, 2) Data Preparation, (3) Model Training, (4) Synthetic Data Generation, and (5) Evaluation. Experimental results on three benchmark datasets show that the proposed work produced data that closely adheres to the statistical distribution of the original dataset, with Wasserstein Distance < 0.05 for numerical features and Jensen-Shannon Divergence < 0.08 for categorical features. Additionally, models trained on datasets including synthetic and real data achieved up to 15% improvement in classification accuracy compared to those trained on real and small datasets alone. Training on a combination of real and synthetic data for the minority class in large datasets significantly improves the F1-score, with gains of approximately 9–10%. This approach also yields a modest increase in overall accuracy (around 1.5%), suggesting enhanced model generalisation. These results indicate that the adapted CTGAN is a viable option for data augmentation, addressing problems with limited and imbalanced data for machine learning data training.
References
Ahsan, M., Gomes, R., Denton, A., Alom, M. Z., Mohammad, N., & Mahmood, M. (2025). Hybrid oversampling technique for imbalanced pattern recognition: Enhancing performance with borderline synthetic minority oversampling and generative adversarial networks (BSGAN). Machine Learning with Applications, 20, Article 100637. https://doi.org/10.1016/j.mlwa.2024. 100637
Biau, G., Sangnier, M., & Tanielian, U. (2021). Some theoretical insights into wasserstein gans. Journal of Machine Learning Research, 22. https://doi.org/10.48550/arXiv.2006.02682
Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16. https://doi.org/10.1613/jair.953
Chen, J., Yang, Z., & Yang, D. (2020). MixText: Linguistically-informed interpolation of hidden space for semi-supervised text classification. Proceedings of the Annual Meeting of the Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.194
Dua, M., Joshi, S., & Dua, S. (2023). Data augmentation based novel approach to automatic speaker verification system. E-Prime - Advances in Electrical Engineering, Electronics and Energy, 6. https://doi.org/10.1016/j.prime.2023.100346
Dunmore, A., Jang-Jaccard, J., Sabrina, F., & Kwak, J. (2023). A comprehensive survey of generative adversarial networks (GANs) in cybersecurity intrusion detection. IEEE Access, 11. https://doi.org/10.1109/ACCESS.2023.3296707
Edunov, S., Ott, M., Auli, M., & Grangier, D. (2018). Understanding back-translation at scale. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,
EMNLP 2018. https://doi.org/10.18653/v1/d18-1045
Fernández, A., García, S., Galar, M., Prati, R. C., Krawczyk, B., & Herrera, F. (2018). Learning from Imbalanced Data Sets. Learning from Imbalanced Data Sets. https://doi.org/10.1007/978-3319-98074-
Fidon, L., Ourselin, S., & Vercauteren, T. (2021). Generalized wasserstein dice score, distributionally robust deep learning, and ranger for brain tumor segmentation: BraTS 2020 challenge. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 12659 LNCS. https://doi.org/10.1007/978-3-030-720872_18
Goodfellow et al. (2020). Generative adversarial networks. Communications of the ACM, 63(11). https://doi.org/10.1145/3422622
Gui, J., Sun, Z., Wen, Y., Tao, D., & Ye, J. (2023). A review on generative adversarial networks: algorithms, theory, and applications. IEEE Transactions on Knowledge and Data Engineering, 35(4). https://doi.org/10.1109/TKDE.2021.3130191
Habibi, O., Chemmakha, M., & Lazaar, M. (2023). Imbalanced tabular data modelization using CTGAN and machine learning to improve IoT Botnet attacks detection. Engineering Applications of Artificial Intelligence, 118. https://doi.org/10.1016/j.engappai.2022.105669
Hou, L., Cao, Q., Shen, H., Pan, S., Li, X., & Cheng, X. (2022). Conditional GANs with auxiliary discriminative classifier. Proceedings of Machine Learning Research, 162. https://doi.org/10.48550/arXiv.1610.09585
Iglesias, G., Talavera, E., González-Prieto, Á., Mozo, A., & Gómez-Canaval, S. (2023). Data Augmentation techniques in time series domain: A survey and taxonomy. Neural Computing and Applications, 35(14). https://doi.org/10.1007/s00521-023-08459-
Isensee, F., Jäger, P. F., Full, P. M., Vollmuth, P., & Maier-Hein, K. H. (2021). nnU-Net for brain tumor segmentation. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 12659 LNCS. https://doi.org/10.1007/978-3-030-72087-2_
Khan, A. A., Chaudhari, O., & Chandra, R. (2024). A review of ensemble learning and data augmentation models for class imbalanced problems: Combination, implementation and evaluation. Expert Systems with Applications, 244. https://doi.org/10.1016/j.eswa.2023.122778
Khan, A. R., Khan, S., Harouni, M., Abbasi, R., Iqbal, S., & Mehmood, Z. (2021). Brain tumor segmentation using K-means clustering and deep learning with synthetic data augmentation for classification. Microscopy Research and Technique, 84(7). https://doi.org/10.1002/jemt.23694
Kingma, D. P., & Welling, M. (2019). An introduction to variational autoencoders. Foundations and Trends in Machine Learning, 12(4). https://doi.org/10.1561/2200000056
Kobayashi, S. (2018). Contextual augmentation: Data augmentation bywords with paradigmatic relations. NAACL HLT 2018-2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, 2. https://doi.org/10.18653/v1/n18-2072
Lyu, Y., & Tian, X. (2025). MWG-UNet++: Hybrid transformer U-Net model for brain tumor segmentation in MRI scans. Bioengineering, 12(2), 140. https://doi.org/10.3390/bioengineering 12020140
Ma, H., Geng, M., Wang, F., Zheng, W., Ai, Y., & Zhang, W. (2024). Data augmentation of a corrosion dataset for defect growth prediction of pipelines using conditional tabular generative adversarial networks. Materials, 17(5). https://doi.org/10.3390/ma17051142
Majeed, A., & Hwang, S. O. (2023). CTGAN-MOS: Conditional generative adversarial network based minority-class-augmented oversampling scheme for imbalanced problems. IEEE Access, 11. https://doi.org/10.1109/ACCESS.2023.3303509
Park et al. (2019). Specaugment: A simple data augmentation method for automatic speech recognition. Proceedings of the Annual Conference of the International Speech Communication
Association, INTERSPEECH, 2019-September. https://doi.org/10.21437/Interspeech.20192680
Patki, N., Wedge, R., & Veeramachaneni, K. (2016). The synthetic data vault. Proceedings - 3rd IEEE International Conference on Data Science and Advanced Analytics, DSAA 2016. https://doi.org/10.1109/DSAA.2016.49
Qiu, H., Yu, B., Gong, D., Li, Z., Liu, W., & Tao, D. (2021). SynFace: Face recognition with synthetic data. Proceedings of the IEEE International Conference on Computer Vision. https://doi.org/10.1109/ICCV48922.2021.01070
Sainin, M. S., Alfred, R., & Ahmad, F. (2021). Ensemble meta classifier with sampling and feature selection for data with imbalance multiclass problem. Journal of Information and Communication Technology, 20(2), 103–133. https://doi.org/10.32890/jict2021.20.2.
Sennrich, R., Haddow, B., & Birch, A. (2016). Improving neural machine translation models with monolingual data. 54th Annual Meeting of the Association for Computational Linguistics,
ACL 2016 - Long Papers, 1. https://doi.org/10.18653/v1/p16-1009
Strelcenia, E., & Prakoonwit, S. (2023). A survey on GAN techniques for data augmentation to address the imbalanced data issues in credit card fraud detection. Machine Learning and Knowledge Extraction, 5(1). https://doi.org/10.3390/make5010019
Tavor et al. (2020). Do not have enough data? Deep learning to the rescue! AAAI 2020-34th AAAI Conference on Artificial Intelligence. https://doi.org/10.1609/aaai.v34i05.6233
Tsalera, E., Papadakis, A., Pagiatakis, G., & Samarakou, M. (2025). Impact evaluation of sound dataset augmentation and synthetic generation upon classification accuracy. Journal of Sensor and Actuator Networks, 14(5), 91. https://doi.org/10.3390/jsan14050091
Uchitomi, H., Ming, X., Zhao, C., Ogata, T., & Miyake, Y. (2023). Classification of mild Parkinson’s disease: Data augmentation of time-series gait data obtained via inertial measurement units. Scientific Reports, 13(1). https://doi.org/10.1038/s41598-023-39862-
Wei, J., & Zou, K. (2019). EDA: Easy data augmentation techniques for boosting performance on text classification tasks. Proceedings of the EMNLP-IJCNLP 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing. https://doi.org/10.18653/v1/d19-1670
Wongvorachan, T., He, S., & Bulut, O. (2023). A comparison of undersampling, oversampling, and smote methods for dealing with imbalanced classification in educational data mining. Information (Switzerland), 14(1). https://doi.org/10.3390/info14010054
Xu, L., Skoularidou, M., Cuesta-Infante, A., & Veeramachaneni, K. (2019). Modeling tabular data using conditional GAN. Advances in Neural Information Processing Systems, 32. https://doi.org/10.48550/arXiv.1907.00503
Zarandah, Q. M. M., Daud, S. M., & Abu-Naser, S. S. (2023). Spectrogram flipping: A new technique for audio augmentation. Journal of Theoretical and Applied Information Technology, 101(11). http://www.jatit.org/volumes/Vol101No11/26Vol101No11.pdf
Zhang et al. (2023). GAN-based one-dimensional medical data augmentation. Soft Computing, 27(15). https://doi.org/10.1007/s00500-023-08345-z
Zhao, X., & Guan, S. (2023). CTCN: A novel credit card fraud detection method based on conditional tabular generative adversarial networks and temporal convolutional network. PeerJ Computer Science, 9. https://doi.org/10.7717/PEERJ-CS.1634
Published
Issue
Section
License
Copyright (c) 2026 Journal of Information and Communication Technology

This work is licensed under a Creative Commons Attribution 4.0 International License.
How to Cite
Research impact
Harvested 2026-09-22Counts differ between services because each indexes a different body of literature. None of them is the whole picture.
2002 - 2020






















