Adaptive Variable Extractions with LDA for Classification of Mixed Variables, and Applications to Medical Data

Authors

  • Hashibah Hamid School of Quantitative Sciences, Universiti Utara Malaysia, Malaysia
  • Nor Idayu Mahat School of Quantitative Sciences, Universiti Utara Malaysia, Malaysia
  • Safwati Ibrahim Institute of Engineering Mathematics, Universiti Malaysia Perlis, Malaysia

DOI:

https://doi.org/10.32890/jict2021.20.3.2

Keywords:

Classification, linear discriminant analysis, multiple correspondence analysis, mixed variables, principal component analysis

Abstract

The strategy surrounding the extraction of a number of mixed variables is examined in this paper in building a model for Linear Discriminant Analysis (LDA). Two methods for extracting crucial variables from a dataset with categorical and continuous variables were employed, namely, multiple correspondence analysis (MCA) and principal component analysis (PCA). However, in this case, direct use of either MCA or PCA on mixed variables is impossible due to restrictions on the structure of data that each method could handle. Therefore, this paper executes some adjustments including a strategy for managing mixed variables so that those mixed variables are equivalent in values. With this, both MCA and PCA can be performed on mixed variables simultaneously. The variables following this strategy of extraction were then utilised in the construction of the LDA model before applying them to classify objects going forward. The suggested models, using three real sets of medical data were then tested, where the results indicated that using a combination of the two methods of MCA and PCA for extraction and LDA could reduce the model’s size, having a positive effect on classifying and better performance of the model since it leads towards minimising the leave-one-out error rate. Accordingly, the models proposed in this paper, including the strategy that was adapted was successful in presenting good results over the full LDA model. Regarding the indicators that were used to extract and to retain the variables in the model, cumulative variance explained (CVE), eigenvalue, and a non-significant shift in the CVE (constant change), could be considered a useful reference or guideline for practitioners experiencing similar issues in future.

References

Alheety, M. (2020). New versions of liu-type estimator in weighted and non-weighted mixed regression model. Baghdad Science Journal, 17(1(Suppl.), 0361. http://bsj.uobaghdad.edu.iq/ index.php/BSJ/article/view/ 5022

Ali, F., Dissanayake, D., Bell, M., & Farrow, M. (2018). Investigating car users’ attitudes to climate change using multiple correspondence analysis. Journal of Transport Geography, 72, 237-247. https://doi.org/10.1016/j.jtrangeo.2018.09.007

AL-Jumaili, A. A. (2020). Hybrid method of linguistic and statistical features for Arabic sentiment analysis. Baghdad Science Journal, 17(1(Suppl.), 0385. https://doi.org/10.21123/ bsj.2020.17.1(Suppl.).0385 Journal of ICT, 20, No. 3 (July) 2021, pp: 305–

An, J., & Chen, Y. P. P. (2009). Finding rule groups to classify high dimensional gene expression datasets. Computational Biology and Chemistry, 33, 108-113. https://doi.org/10.1016/j. compbiolchem.2008.07.031

Anderson, T. W. (1958). An introduction to multivariate statistical analysis. New York: John Wiley & Sons, Inc.

Artoni, F., Delorme, A., & Makeig, S. (2018). Applying dimension reduction to EEG data by principal component analysis reduces the quality of its subsequent independent component decomposition. Neurolmage, 175, 176-187. https://doi.org/10.1016/j.neuroimage.2018.03.016

Barnouti, N. H., Al-Dabbagh, S. S., Matti, W. E., & Naser, M. A. (2016). Face detection and recognition using Viola-Jones with PCA-LDA and square Euclidean distance. International Journal of Advanced Computer Science and Applications (IJACSA), 7(5), 371-377. https://doi.org/10.14569/ijacsa.2016.070550

Blasius, J., & Thiessen, V. (2000). Methodological artifacts in measures of political efficacy and trust: A multiple correspondence analysis. Political Analysis, 9(1), 1-20. https://doi.org/10.1093/ oxfordjournals.pan.a004862

Bodnar, T., Mazur, S., Ngailo, E., & Parolya, N. (2020). Discriminant analysis in small and large dimensions. Theory of Probability and Mathematical Statistics, (100), 21-41. https://doi.org/10.1090/tpms/1096

Chandan, M., White, H., & Wuyts, M. (1998). Econometrics and data analysis for developing countries. London: Routledge.

Das, S., & Sun, X. (2016). Association knowledge for fatal run-off-road crashes by multiple correspondence analysis. IATSS Research, 39(2), 146-155. https://doi.org/10.1016/j.iatssr.2015.07.001

Das, S., Avelar, R., Dixon, K., & Sun, X. (2018). Investigation on the wrong way driving crash patterns using multiple correspondence analysis. Accident Analysis & Prevention, 111, 43-55. https://doi.org/10.1016/j.aap.2017.11.016

D’Enza, A. I., & Greenacre, M. J. (2012). Multiple correspondence analysis for the quantification and visualization of large categorical data sets. In A. Di Ciaccio, M. Coli & J. M. A. Ibaňez (Eds.), Advanced statistical methods for the analysis of large data-sets: Studies in theoretical and applied statistics (pp. 453-463). Berlin Heidelberg: Springer-Verlag. https://doi.org/10.1007/978-3-642-21037-2_41

Deshpande, N. T., & Ravishhankar, S. (2017). Face detection and recognition using Viola-Jones algorithm and fusion of PCA and ANN. Advances in Computational Sciences and Technology, 10(5), 1173-1189. Journal of ICT, 20, No. 3 (July) 2021, pp: 305–

Dungey, M., Tchatoka, F. D., & Yanotti, M. B. (2018). Using multiple correspondence analysis for finance: A tool for assessing financial inclusion. International Review of Financial Analysis, 59, 212-222. https://doi.org/10.1016/j.irfa.2018.08.007

Ghosh, J., & Shuvo, S. B. (2019). Improving classification model’s performance using linear discriminant analysis on linear data. In the 10th International Conference on Computing, Communica-tion and Networking Technologies (ICCCNT) 2019 July 6 (pp. 1-5). IEEE. https://doi.org/10.1109/icccnt45670.2019.8944632

Guttman, L. (1941). The quantification of a class of attributes: A theory and method of scale construction. In P. Horst, P. Wallin & L. Guttman (Eds.), The prediction of personal adjustment (pp. 319-348). New York: Social Science Research Council.

Gyamfi, K. S., Brusey, J., Hunt, A., & Gaura, E. (2017). Linear classifier design under heteroscedasticity in linear discriminant analysis. Expert Systems with Applications, 79, 44-52. https://doi.org/10.1016/j.eswa.2017.02.039

Hamid, H., Mei, L. M., & Yahaya, S. S. S. (2017). New discrimination procedure of location model for handling large categorical variables. Sains Malaysiana, 46(6), 1001-1010. https://doi.org/10.17576/jsm-2017-4606-20

Hamid, H., Ngu, P. A., & Alipiah, F. M. (2018). New smoothed location models integrated with PCA and two types of MCA for handling large number of mixed continuous and binary variables. Pertanika Journal of Science & Technology, 26(1), 247-260.

Jamal, A., Handayani, A., Septiandiri, A. A., Ripmiatin, E., & Effendi, Y. (2018). Dimensionality reduction using PCA and K-means clustering for breast cancer prediction. Lontar Komputer: Jurnal Ilmiah Teknologi Imformasi, 192-201. https://doi.org/10.24843/lkjiti.2018.v09.i03.p08

Jolliffe, I. T. (1986). Principal component analysis. New York: Springer-Verlag.

Josse, J., & Husson, F. (2016). missMDA: A package for handling missing values in multivariate data analysis. Journal of Statistical Software, 70(1), 1-31. https://doi.org/10.18637/jss.v070.i01

Kaiser, H. F. (1960). The application of electronic computers to factor analysis. Educational and Psychological Measurement, 20, 141-151. https://doi.org/10.1177/001316446002000116

Kaminska, A., Ickowicz, A., Plouin, P., Bru, M. F., Dellatolas, G., & Dulac, O. (1999). Delineation of cryptogenic lennox–gastaut syndrome and myoclonic astatic epilepsy using multiple correspondence analysis. Epilepsy Research, 36, 15-29. https://doi.org/10.1016/s0920-1211(99)00021-2 Journal of ICT, 20, No. 3 (July) 2021, pp: 305–

Krzanowski, W. J. (1975). Discrimination and classification using both binary and continuous variables. Journal of the American Statistical Association, 70(352), 782-790. https://doi.org/10.10 80/01621459.1975.10480303

Krzanowski, W. J. (1977). The performance of fisher’s linear discriminant function under non-optimal conditions. Technometrics, 19, 191-200. https://doi.org/10.1080/0040170 6.1977.10489527

Krzanowski, W. J. (1980). Mixtures of continuous and categorical variables in discriminant analysis. Biometrics, 36, 493-499. https://doi.org/10.2307/2530217

Krzanowski, W. J. (1987). Selection of variables to preserve multivariate data structure using principal components. Applied Statistics, 36(1), 22-33. https://doi.org/10.2307/2347842

Li, M. (2017). Application of cart decision tree combined with PCA algorithm in intrusion detection. In the 8th International Conference on Software Engineering and Service Science (ICSESS) Nov 24 (pp. 38-41). IEEE. https://doi.org/10.1109/ icsess.2017.8342859

Mahat, N. I., Krzanowski, W. J., & Hernandez, A. (2007). Variable selection in discriminant analysis based on the location model for mixed variables. Advances in Data Analysis and Classification, 1(2), 105-122. https://doi.org/10.1007/s11634-007-0009-9

Messaoud, R. B., Boussaid, O., & Rabaséda, S. L. (2007). A multiple correspondence analysis to organize data cubes. Databases and Information Systems IV: Frontiers in Artificial Intelligence and Applications, 155(1), 133-146.

Mohamed, R., Zainudin, M. N. S., Sulaiman, M. N., Perumal, T., & Mustapha, N. (2018). Multi-label classification for physical activity recognition from various accelerometer sensor positions. Journal of Information and Communication Technology, 17(2), 209–231. https://doi.org/10.32890/jict2018.17.2.3

Mori, Y., Kuroda, M., & Makino, N. (2016). Sparse multiple correspondence analysis: Nonlinear principal component analysis and its applications. Singapore: Springer. https://doi.org/10.1007/978-981-10-0159-8_5

Nasution, M. Z. F., Sitompul, O. S., & Ramli, M. (2018). PCA based feature reduction to improve the accuracy of decision tree C4.5 classification. Journal of Physics: IOP Conference Series, 978, 012058. https://doi.org/10.1088/1742-6596/978/1/012058

Nazman, E., & Erbas, S. (2017). Evaluation of group homogeneity in Gaussian mixture models using combined cluster and discriminant analysis. Sinop University Journal of Natural Sciences, 2(1), 121-132. Journal of ICT, 20, No. 3 (July) 2021, pp: 305–

Peres, F. A., & Fogliatto, F. S. (2018). Variable selection methods in multivariate statistical process control: A systematic literature review. Computers & Industrial Engineering, 115, 603-619. https://doi.org/10.1016/j.cie.2017.12.006

Prats-Montalbán, J. M., Ferrer, A., Malo, J. L., & Gorbeña, J. A. (2006). Comparison of different discriminant analysis techniques in a steel industry welding process. Chemometrics and Intelligent Laboratory Systems, 80, 109-119. https://doi.org/10.1016/j. chemolab.2005.08.005

Sivasankaran, S. K., & Balasubramanian, V. (2020). Investigation of pedestrian crashes using multiple correspondence analysis in India. International Journal of Injury Control and Safety Promotion, 27(2), 144-155. https://doi.org/10.1080/17457300.2019.1681005

Stevens, J. P. (2002). Applied multivariate statistics for the social sciences (4th edition). New Jersey: Lawrence Erlbaum Associates, Inc.

Swesi, I. M. A. O., & Bakar, A. A. (2019). Feature clustering for PSO-based feature construction on high-dimensional data. Journal of Information and Communication Technology, 18(4), 439-472. https://doi.org/10.32890/jict2019.18.4.3

Tarr, G., Müller, S., & Weber, N. C. (2016). Robust estimation of precision matrices under cellwise contamination. Computational Statistics & Data Analysis, 93, 404-420. https://doi.org/10.1016/j.csda.2015.02.005

Tharwat, A., Gaber, T., Ibrahim, A., & Hassanien, A. E. (2017). Linear discriminant analysis: A detailed tutorial. AI Communications, 30(2), 169-190. https://doi.org/10.3233/aic-170729

Vyas, S., & Kumaranayake, L. (2006). Constructing socio-economic status indices: How to use principal components analysis. Health Policy and Planning, 21(6), 459-468. https://doi.org/10.1093/heapol/czl029

Zhang, D. F., Chen, Y. C., Chen, H., Zhang, W. D., Sun, J., Mao, C. N., Su, W., Wang, P., & Yin, X. (2017). A high-resolution MRI study of relationship between remodelling patterns and ischemic stroke in patients with atherosclerotic middle cerebral artery stenosis. Frontiers in Aging Neuroscience, 9, 140. https://doi.org/10.3389/fnagi.2017.00140

Downloads

Published

11-06-2021

How to Cite

Hamid, H., Mahat, N. I., & Ibrahim, S. (2021). Adaptive Variable Extractions with LDA for Classification of Mixed Variables, and Applications to Medical Data. Journal of Information and Communication Technology, 20(3), 305-327. https://doi.org/10.32890/jict2021.20.3.2

Research impact

Harvested 2026-09-06
0 citations recorded so far

Counts differ between services because each indexes a different body of literature. None of them is the whole picture.

Identifiers DOI 10.32890/jict2021.20.3.2 OpenAlex W3175530666 Scopus 85109446905

Most read articles by the same author(s)