Enhancing the Effectiveness of Machine Learning-Based Phishing Email Detection via an Improved Pre-Processing Technique for Data Security
DOI:
https://doi.org/10.32890/jict2025.24.4.4Keywords:
Data protection, machine learning, phishing email detection, pre-processing, supervised learningAbstract
The growing volume and sophistication of phishing emails have become a significant threat to data security, often serving as the initial vector for data breaches. Most past studies have focused on comparing machine learning models to determine the best-performing algorithm. They often neglect the role of pre-processing, which also contributes to the effectiveness of these models. To address this gap, this study investigates the impact of pre-processing techniques on phishing email detection, aiming to strengthen data protection. Three supervised machine learning algorithms, which are Support Vector Machine (SVM), Random Forest and Decision Tree, were selected to undergo two experimental iterations: one with basic pre-processing and the other with an enhanced pre-processing technique including Synthetic Minority Oversampling Technique (SMOTE), Term Frequency-Inverse Document Frequency (TF-IDF), Singular Value Decomposition (SVD) and cross-validation. Using a dataset comprising 28,747 labelled emails, the models were trained, tested, and evaluated based on accuracy, precision, recall, and F1-score, with further insight gained through confusion matrix analysis. Among the models, Random Forest demonstrated the strongest consistent performance across all metrics, while Decision Tree showed the most notable improvement. Although SVM maintains high recall and precision, it is less responsive to the applied pre-processing techniques. This result demonstrates that pre-processing techniques significantly contribute to the performance of the detection models. Overall, these findings highlight the critical role of pre-processing in enhancing phishing email detection, which contributes to stronger organisational resilience.
References
Ahmed, M. R., Islam, M. M., & Layek, M. A. (2024). Phishing URL detection using comprehensive feature extraction and machine learning techniques. IEEE CS BDC Symposium 2024. https://s24.ieeecsbdc.org/papers/156
Alkhalil, Z., Hewage, C., Nawaf, L., & Khan, I. (2021). Phishing attacks: A recent comprehensive study and a new anatomy. Frontiers in Computer Science, 3, 1-23. https://doi.org/10.3389/fcomp. 2021.563060
Al-Yozbaky, R. S., & Alanezi, M. (2023). Phishing emails detection models: A comparative study. Journal of Modern Computing and Engineering Research, 2023, 58–67. https://jmcer.org/ research/phishing-emails-detection-models-a-comparative-study/
Anti-Phishing Working Group (APWG). (2024). Phishing activity trends report (4th quarter 2024). https://docs.apwg.org/reports/apwg_trends_report_q4_2024.pdf
Chy, M. K. H. (2024). Securing the web: Machine learning’s role in predicting and preventing phishing attacks. International Journal of Science and Research Archive, 13(1), 1004-1011. https://doi.org/10.30574/ijsra.2024.13.1.1770
Divakaran, D. M., & Oest, A. (2022). Phishing detection leveraging machine learning and deep learning: A review. arXiv. https://doi.org/10.48550/arXiv.2205.07411
Felix, E. A. & Lee, S. P. (2019). Systematic literature review of pre-processing techniques for imbalance data. IET Software, 13(6), 479-496. https://doi.10.1049/iet-sen 2018.5193
Gualberto, E. S., De Sousa, R. T., Vieira, T. P. D. B., Da Costa, J. P. C. L., & Duque, C. G. (2020). From feature engineering and topics models to enhance prediction rates in phishing detection. IEEE Access, 8, 76368-76385. https://doi.org/10.1109/ACCESS.2020.2989126
Habib, P., Sharma, U., & Sethi, K. S. (2022). Phishing detection with machine learning. International Journal for Study in Applied Science & Engineering Technology (IJRASET), 10(12), 16091615. https://doi.org/10.22214/ijraset.2022.48276
Hamadouche, S., Boudraa, O., & Gasmi, M. (2024). Combining lexical, host, and content-based features for phishing websites detection using machine learning models. EAI Endorsed Transactions on Scalable Information Systems, 11(6), 1-15. https://doi.org/10.4108/eetsis.4421
Harikrishnan, R., Soman, K. P., & Jayaraj, V. (2018). A machine learning approach towards phishing email detection CEN-Security@IWSPA 2018. Proceedings of the 1st Anti-Phishing Shared Task Pilot at 4th ACM IWSPA co-located with 8th ACM Confer-ence on Data and Application
Security and Privacy (CODASPY 2018) (pp. 21-28). Tempe, Arizona, USA: CEUR-WS.org. https://ceur-ws.org/Vol-2124/paper_7.pdf
Hina, M., Ali, M., Javed, A. R., Srivastava, G., Gadekallu, T. R., & Jalil, Z. (2021). Email classification and forensics analysis using machine learning. 2021 IEEE Smart-World, Ubiquitous Intelligence & Computing, Advanced & Trusted Computing, Scalable Computing & Communications, Internet of People and Smart City Innovation (SmartWorld/SCALCOM/ UIC/ATC/IOP/SCI) (pp. 630–635). IEEE. https://doi.org/10.1109/SWC50871.2021.00093
Hussin, I. H., Nazarudin, L. S. I., & Nor Azman, N. A. (2024). Comparative analysis of machine learning and deep learning for email spam detection. TechRxiv. https://doi.org/10.36227/ techrxiv.172115119.92836191/v1
Johnson, J. M., & Khoshgoftaar, T. M. (2019). Survey on deep learning with class imbalance. Journal of Big Data, 6(1), 1-54. https://doi.org/10.1186/s40537-019-0192-5
Konda, B. (2022). The impact of data pre-processing on data mining outcomes. World Journal of
Advanced Research and Reviews, 15(3), 540–544. https://doi.org/10.30574/wjarr.2022.15.3. 0931
Kovacs, E. (2024). Discount retail giant pepco loses €15 million to cybercriminals. SecurityWeek. https://www.securityweek.com/discount-retail-giant-pepco-loses-e15-million-to-cybercrimina ls/
Kumari, A., Punn, N. S., Sonbhadra, S. K., & Agarwal, S. (2023). Impact of the composition of feature extraction and class sampling in Medicare fraud detection. In M. Tanveer, S. Agarwal, S.
Ozawa, A. Ekbal, & A. Jatowt (Eds.), Neural information processing. ICONIP 2022 (Lecture Notes in Computer Science, Vol. 13625, pp. 708–720). Springer. https://doi.org/10.1007/9783-031-30111-7_54
Malik, H. K., Al-Anber, N. J., & Al-Mekhlafi, F. A. E. (2023). Comparison of feature selection and feature extraction role in dimensionality reduction of big data. Journal of Techniques, 5(1), 184–192. https://doi.org/10.51173/jt.v5i1.1027
Markoulidakis, I. & Markoulidakis, G. (2024). Probabilistic confusion matrix: A novel method for machine learning algorithm generalised performance analysis. Technologies, 12(7), 113. https://doi.org/10.3390/technologies12070113
Mohey, H. & Mohsen, S. (2021). Using machine learning techniques for predicting email spam. Special Issue The Fifth Scientific Conference of Obour Institutes “Engineering and Informatics for Sustainable Smart Organizations (pp. 19-23). Obour Institutes, Egypt: International Journal of Instructional Technology and Educational Studies (IJITES). Reddy, G. D., Sreelata, S., & Shreya, D. (Year). Efficient email phishing detection using machine learning. International Journal of Scientific Study in Computer Science, Engineering and Information Technology, 9(3), 356-360. https://doi.org/10.32628/IJSRCSEIT
Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5), 206–215. https://doi.org/10.1038/s42256-019-0048-x
Shejole, S., Salam, T. A., Gupta, M. K., Sharma, K., & Attar, S. (2023). Email spam detection using machine learning. International Study Journal of Engineering and Technology (IRJET), 10(5). https://www.irjet.net/archives/V10/i5/IRJET-V10I5226.pdf
Tanti, R. (2024). Phishing attack: A case study and their prevention techniques. International Journal for Multidisciplinary Research (IJFMR), 6(5), 1–8. https://doi.org/10.55041/IJSREM38042
Tawil, A. A., Almazaydeh, L., Qawasmeh, D., Qawasmeh, B., Alshinwan, M. & Elleithy, K. (2024). Comparative analysis of machine learning algorithms for email phishing detection using TFIDF, Word2Vec, and BERT. Computers, Materials & Continua, 81(2), 3395–3412. https://doi.org/10.32604/cmc.2024.057279
The Radicati Group, Inc. (2023). Email statistics report, 2023-2027. The Radicati Group, Inc. https://www.radicati.com/wp/wp-content/uploads/2023/04/Email-Statistics-Report-2023-2027-ExecutiveSummary.pdf
Published
Issue
Section
License
Copyright (c) 2026 Journal of Information and Communication Technology

This work is licensed under a Creative Commons Attribution 4.0 International License.
How to Cite
Research impact
Harvested 2026-09-05Counts differ between services because each indexes a different body of literature. None of them is the whole picture.
2002 - 2020






















