Conditional Tabular Generative Adversarial Network-based Synthetic Data Generation for Model Generalisation Improvement

Authors

  • Yuhanis Yusof School of Computing, Universiti Utara Malaysia, Malaysia
  • Fathima Fajila School of Computing, Universiti Utara Malaysia, Faculty of Applied Sciences, South Eastern University of Sri Lanka, Malaysia

DOI:

https://doi.org/10.32890/jict2026.25.1.1

Keywords:

Deep learning, CTGAN, data augmentation, synthetic data

Abstract

Accessing extensive and varied datasets is essential for developing strong predictive models in data analytics. However, many real-world applications suffer from small and imbalanced datasets, leading to overfitting, poor generalisation, and low model performance. Traditional data augmentation techniques are often unsuitable for tabular data, as they fail to preserve complex feature relationships. To address this challenge, this study adapts the Conditional Tabular Generative Adversarial Network (CTGAN) for synthetic data generation. The proposed approach involves five phases: (1) Data Acquisition, 2) Data Preparation, (3) Model Training, (4) Synthetic Data Generation, and (5) Evaluation.  Experimental results on three benchmark datasets show that the proposed work produced data that closely adheres to the statistical distribution of the original dataset, with Wasserstein Distance < 0.05 for numerical features and Jensen-Shannon Divergence < 0.08 for categorical features. Additionally, models trained on datasets including synthetic and real data achieved up to 15% improvement in classification accuracy compared to those trained on real and small datasets alone. Training on a combination of real and synthetic data for the minority class in large datasets significantly improves the F1-score, with gains of approximately 9–10%. This approach also yields a modest increase in overall accuracy (around 1.5%), suggesting enhanced model generalisation. These results indicate that the adapted CTGAN is a viable option for data augmentation, addressing problems with limited and imbalanced data for machine learning data training.

References

Ahsan, M., Gomes, R., Denton, A., Alom, M. Z., Mohammad, N., & Mahmood, M. (2025). Hybrid oversampling technique for imbalanced pattern recognition: Enhancing performance with borderline synthetic minority oversampling and generative adversarial networks (BSGAN). Machine Learning with Applications, 20, Article 100637. https://doi.org/10.1016/j.mlwa.2024. 100637

Biau, G., Sangnier, M., & Tanielian, U. (2021). Some theoretical insights into wasserstein gans. Journal of Machine Learning Research, 22. https://doi.org/10.48550/arXiv.2006.02682

Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16. https://doi.org/10.1613/jair.953

Chen, J., Yang, Z., & Yang, D. (2020). MixText: Linguistically-informed interpolation of hidden space for semi-supervised text classification. Proceedings of the Annual Meeting of the Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.194

Dua, M., Joshi, S., & Dua, S. (2023). Data augmentation based novel approach to automatic speaker verification system. E-Prime - Advances in Electrical Engineering, Electronics and Energy, 6. https://doi.org/10.1016/j.prime.2023.100346

Dunmore, A., Jang-Jaccard, J., Sabrina, F., & Kwak, J. (2023). A comprehensive survey of generative adversarial networks (GANs) in cybersecurity intrusion detection. IEEE Access, 11. https://doi.org/10.1109/ACCESS.2023.3296707

Edunov, S., Ott, M., Auli, M., & Grangier, D. (2018). Understanding back-translation at scale. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,

EMNLP 2018. https://doi.org/10.18653/v1/d18-1045

Fernández, A., García, S., Galar, M., Prati, R. C., Krawczyk, B., & Herrera, F. (2018). Learning from Imbalanced Data Sets. Learning from Imbalanced Data Sets. https://doi.org/10.1007/978-3319-98074-

Fidon, L., Ourselin, S., & Vercauteren, T. (2021). Generalized wasserstein dice score, distributionally robust deep learning, and ranger for brain tumor segmentation: BraTS 2020 challenge. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 12659 LNCS. https://doi.org/10.1007/978-3-030-720872_18

Goodfellow et al. (2020). Generative adversarial networks. Communications of the ACM, 63(11). https://doi.org/10.1145/3422622

Gui, J., Sun, Z., Wen, Y., Tao, D., & Ye, J. (2023). A review on generative adversarial networks: algorithms, theory, and applications. IEEE Transactions on Knowledge and Data Engineering, 35(4). https://doi.org/10.1109/TKDE.2021.3130191

Habibi, O., Chemmakha, M., & Lazaar, M. (2023). Imbalanced tabular data modelization using CTGAN and machine learning to improve IoT Botnet attacks detection. Engineering Applications of Artificial Intelligence, 118. https://doi.org/10.1016/j.engappai.2022.105669

Hou, L., Cao, Q., Shen, H., Pan, S., Li, X., & Cheng, X. (2022). Conditional GANs with auxiliary discriminative classifier. Proceedings of Machine Learning Research, 162. https://doi.org/10.48550/arXiv.1610.09585

Iglesias, G., Talavera, E., González-Prieto, Á., Mozo, A., & Gómez-Canaval, S. (2023). Data Augmentation techniques in time series domain: A survey and taxonomy. Neural Computing and Applications, 35(14). https://doi.org/10.1007/s00521-023-08459-

Isensee, F., Jäger, P. F., Full, P. M., Vollmuth, P., & Maier-Hein, K. H. (2021). nnU-Net for brain tumor segmentation. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 12659 LNCS. https://doi.org/10.1007/978-3-030-72087-2_

Khan, A. A., Chaudhari, O., & Chandra, R. (2024). A review of ensemble learning and data augmentation models for class imbalanced problems: Combination, implementation and evaluation. Expert Systems with Applications, 244. https://doi.org/10.1016/j.eswa.2023.122778

Khan, A. R., Khan, S., Harouni, M., Abbasi, R., Iqbal, S., & Mehmood, Z. (2021). Brain tumor segmentation using K-means clustering and deep learning with synthetic data augmentation for classification. Microscopy Research and Technique, 84(7). https://doi.org/10.1002/jemt.23694

Kingma, D. P., & Welling, M. (2019). An introduction to variational autoencoders. Foundations and Trends in Machine Learning, 12(4). https://doi.org/10.1561/2200000056

Kobayashi, S. (2018). Contextual augmentation: Data augmentation bywords with paradigmatic relations. NAACL HLT 2018-2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, 2. https://doi.org/10.18653/v1/n18-2072

Lyu, Y., & Tian, X. (2025). MWG-UNet++: Hybrid transformer U-Net model for brain tumor segmentation in MRI scans. Bioengineering, 12(2), 140. https://doi.org/10.3390/bioengineering 12020140

Ma, H., Geng, M., Wang, F., Zheng, W., Ai, Y., & Zhang, W. (2024). Data augmentation of a corrosion dataset for defect growth prediction of pipelines using conditional tabular generative adversarial networks. Materials, 17(5). https://doi.org/10.3390/ma17051142

Majeed, A., & Hwang, S. O. (2023). CTGAN-MOS: Conditional generative adversarial network based minority-class-augmented oversampling scheme for imbalanced problems. IEEE Access, 11. https://doi.org/10.1109/ACCESS.2023.3303509

Park et al. (2019). Specaugment: A simple data augmentation method for automatic speech recognition. Proceedings of the Annual Conference of the International Speech Communication

Association, INTERSPEECH, 2019-September. https://doi.org/10.21437/Interspeech.20192680

Patki, N., Wedge, R., & Veeramachaneni, K. (2016). The synthetic data vault. Proceedings - 3rd IEEE International Conference on Data Science and Advanced Analytics, DSAA 2016. https://doi.org/10.1109/DSAA.2016.49

Qiu, H., Yu, B., Gong, D., Li, Z., Liu, W., & Tao, D. (2021). SynFace: Face recognition with synthetic data. Proceedings of the IEEE International Conference on Computer Vision. https://doi.org/10.1109/ICCV48922.2021.01070

Sainin, M. S., Alfred, R., & Ahmad, F. (2021). Ensemble meta classifier with sampling and feature selection for data with imbalance multiclass problem. Journal of Information and Communication Technology, 20(2), 103–133. https://doi.org/10.32890/jict2021.20.2.

Sennrich, R., Haddow, B., & Birch, A. (2016). Improving neural machine translation models with monolingual data. 54th Annual Meeting of the Association for Computational Linguistics,

ACL 2016 - Long Papers, 1. https://doi.org/10.18653/v1/p16-1009

Strelcenia, E., & Prakoonwit, S. (2023). A survey on GAN techniques for data augmentation to address the imbalanced data issues in credit card fraud detection. Machine Learning and Knowledge Extraction, 5(1). https://doi.org/10.3390/make5010019

Tavor et al. (2020). Do not have enough data? Deep learning to the rescue! AAAI 2020-34th AAAI Conference on Artificial Intelligence. https://doi.org/10.1609/aaai.v34i05.6233

Tsalera, E., Papadakis, A., Pagiatakis, G., & Samarakou, M. (2025). Impact evaluation of sound dataset augmentation and synthetic generation upon classification accuracy. Journal of Sensor and Actuator Networks, 14(5), 91. https://doi.org/10.3390/jsan14050091

Uchitomi, H., Ming, X., Zhao, C., Ogata, T., & Miyake, Y. (2023). Classification of mild Parkinson’s disease: Data augmentation of time-series gait data obtained via inertial measurement units. Scientific Reports, 13(1). https://doi.org/10.1038/s41598-023-39862-

Wei, J., & Zou, K. (2019). EDA: Easy data augmentation techniques for boosting performance on text classification tasks. Proceedings of the EMNLP-IJCNLP 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing. https://doi.org/10.18653/v1/d19-1670

Wongvorachan, T., He, S., & Bulut, O. (2023). A comparison of undersampling, oversampling, and smote methods for dealing with imbalanced classification in educational data mining. Information (Switzerland), 14(1). https://doi.org/10.3390/info14010054

Xu, L., Skoularidou, M., Cuesta-Infante, A., & Veeramachaneni, K. (2019). Modeling tabular data using conditional GAN. Advances in Neural Information Processing Systems, 32. https://doi.org/10.48550/arXiv.1907.00503

Zarandah, Q. M. M., Daud, S. M., & Abu-Naser, S. S. (2023). Spectrogram flipping: A new technique for audio augmentation. Journal of Theoretical and Applied Information Technology, 101(11). http://www.jatit.org/volumes/Vol101No11/26Vol101No11.pdf

Zhang et al. (2023). GAN-based one-dimensional medical data augmentation. Soft Computing, 27(15). https://doi.org/10.1007/s00500-023-08345-z

Zhao, X., & Guan, S. (2023). CTCN: A novel credit card fraud detection method based on conditional tabular generative adversarial networks and temporal convolutional network. PeerJ Computer Science, 9. https://doi.org/10.7717/PEERJ-CS.1634

Downloads

Published

31-01-2026

How to Cite

Yusof, Y., & Fajila, F. (2026). Conditional Tabular Generative Adversarial Network-based Synthetic Data Generation for Model Generalisation Improvement. Journal of Information and Communication Technology, 25(1), 1-16. https://doi.org/10.32890/jict2026.25.1.1

Research impact

Harvested 2026-09-22
2 citations, from OpenAlex — the highest of the sources checked

Counts differ between services because each indexes a different body of literature. None of them is the whole picture.

Identifiers DOI 10.32890/jict2026.25.1.1 OpenAlex W7128513434 Scopus 105031359437

Most read articles by the same author(s)