Irrelevant Feature and Rule Removal for Structural Associative Classification

Authors

  • Izwan Nizal Mohd Shaharanee School of Quantitative Sciences, Universiti Utara Malaysia, Malaysia
  • Jastini Mohd Jamil School of Quantitative Sciences, Universiti Utara Malaysia, Malaysia

DOI:

https://doi.org/10.32890/jict2015.14.6

Keywords:

Features selection, rules removal, frequent item set mining

Abstract

In the classification task, the presence of irrelevant features can significantly degrade the performance of classification algorithms, in terms of additional processing time, more complex models and the likelihood that the models have poor generalization power due to the over fitting problem. Practical applications of association rule mining often suffer from overwhelming number of rules that are generated, many of which are not interesting or not useful for the application in question. Removing rules comprised of irrelevant features can signifi cantly improve the overall performance. In this paper, we explore and compare the use of a feature selection measure to filter out unnecessary and irrelevant features/attributes prior to association rules generation. The experiments are performed using a number of real-world datasets that represent diverse characteristics of data items. Empirical results confirm that by utilizing feature subset selection prior to association rule generation, a large number of rules with irrelevant features can be eliminated. More importantly, the results reveal that removing rules that hold irrelevant features improve the accuracy rate and capability to retain the rule coverage rate of structural associative association.

 

References

Asuncion, A., & Newman, D. J. (2007). UCI Machine Learning Repository. University of California Irvine School of Information. University of California, Irvine, School of Information and Computer Sciences. Retrieved from http://www.ics.uci.edu/~mlearn/MLRepository.html

Bayardo Jr., R. J., & Agrawal, R. (1999). Mining the Most Interesting Rules. In Proceedings of the Fifth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 145–154). New York, NY, USA: ACM.

Blanchard, J., Guillet, F., Gras, R., & Briand, H. (2005). Using Information-Theoretic Measures to Assess Association Rule Interestingness. In Proceedings of the Fifth IEEE International Conference on Data Mining. IEEE Computer Society.

Bolón-Canedo, V., Sánchez-Maroño, N., & Alonso-Betanzos, A. (2013). A review of feature selection methods on synthetic data. Knowledge and Information Systems, 34(3), 483–519. doi:10.1007/s10115-012-0487-8

Brin, S., Motwani, R., & Silverstein, C. (1997). Beyond market baskets: generalizing association rules to correlations. In Proceedings of the 1997 ACM SIGMOD International Conference on Management of Data. Tucson, Arizona, United States: ACM. Journal of ICT, 14, 2015, pp: 95–

Cheng, H., Yan, X., Han, J., & Hsu, C.-W. (2007). Discriminative frequent pattern analysis for effective classification. In Proceeding of the 23rd International IEEE Conference on Data Engineering (pp. 716-725).

Cheng, H., Yan, X., Han, J., & S. Yu, P. (2008). Direct discriminative pattern mining for effective classification, In Proceeding of the 24th International IEEE Conference on Data Engineering (pp. 67-178).

Dash, M., & Liu, H. (1997). Feature selection for classification. Intelligent Data Analysis, 1(1-4), 131–156. doi:10.1016/S1088-467X(97)00008-5

Geng, L., & Hamilton, H. J. (2006). Interestingness measures for data mining. ACM Computing Surveys, 38(3), 9. doi:10.1145/1132960.1132963

Hadzic, F., & Dillon, T. S. (2006). Using the Symmetrical Tau (t) criterion for feature selection in decision tree and neural network learning. In Proceedings of the SIAM 2nd Workshop on Feature Selection for Data Mining: Interfacing Machine Learning and Statistics.

Han, J., Kamber, M., & Pei, J. (2011). Data mining: Concepts and techniques (3rd ed.). Morgan Kaufmann Publishers.

Hilderman, R., & Hamilton, H. (2001). Evaluation of interestingness measures for ranking discovered knowledge. In Proceedings of the 5th Pacific-Asia Conference on Knowledge Discovery and Data Mining: Springer-Verlag.

Jaroszewicz, S., & Simovici, D. (2001). A general measure of rule interestingness. In L. Raedt & A. Siebes (Eds.), Principles of data mining and knowledge discovery SE-21 (Vol. 2168, pp. 253–265). Springer Berlin Heidelberg.

Ke, Y., Cheng, J., & Ng, W. (2008). An information-theoretic approach to quantitative association rule mining. Knowledge and Information Systems, 16(2), 213–244. doi:10.1007/s10115-007-0104-4

Kudo, M., & Sklansky, J. (2000). Comparison of algorithms that select features for pattern classifiers. Pattern Recognition. doi:10.1016/ S0031-3203(99)00041-2

Li, J., Shen, H., & Topor, R. (2002). Mining the optimal class association rule set. Knowledge-Based Systems, 15, 399–405. doi:10.1016/S0950-7051(02)00024-2

Molina, L. C., Belanche, L., & Nebot, A. (2002). Feature selection algorithms: A survey and experimental evaluation. In Proceding of IEEE International Conference on Data Mining (ICDM ’02) (pp. 306–313).

Olanweraju, R. F., Aburas, A. A., Omran, O. K., & Abdalla, A.-H. H. (2010). Damageless Digital Watermarking using Complex Valued Artificial Neural Network. Journal of Information and Communication Technology, 9, 111–137. Journal of ICT, 14, 2015, pp: 95–

Peng, H., Long, F., & Ding, C. (2005). Feature selection based on mutual information: Criteria of max-dependency, max-relevance, and min-redundancy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 27, 1226–1238. doi:10.1109/TPAMI.2005.159

Quinlan, J. R. (1986). Induction of decision trees. Machine Learning, 1(1), 81–106.

Quinlan, J. R. (1993). C4. 5: Programs for machine learning (Vol. 1). Morgan kaufmann.

Roiger, R. J., & Geatz, M. W. (2003). Data Mining: A Tutorial-Based Primer. Addison Wesley.

Shaharanee, I. N. M., & Hadzic, F. (2013). Evaluation and optimization of frequent, closed and maximal association rule based classification. Statistics and Computing. doi:10.1007/s11222-013-9404-6

Shaharanee, I. N. M., Hadzic, F., & Dillon, T. S. (2011). Interestingness measures for association rules based on statistical validity. Knowledge-Based Systems, 24(3), 386–392. doi:10.1016/j.knosys. 2010.11.005

Shaharanee, I. N. M., & Jamil, J. M. (2013). Features selection and rule Removal for frequent Association rule based classification. In Proceedings of the 4th International Conference on Computing and Informatic (pp. 377–382).

Smyth, P., & Goodman, R. M. (1992). An information theoretic approach to rule induction from databases. IEEE Transactions on Knowledge and Data Engineering, 4(4), 301–316. doi:10.1109/69.149926

Tan, P. N., Kumar, V., & Srivastava, J. (2002). Selecting the right interestingness measure for association patterns. In Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Edmonton, Alberta, Canada: ACM.

Tan, P.-N., Steinbach, M., & Kumar, V. (2014). Introduction to data mining. Essex: Pearson Education Limited.

Yusof, S. A. M., Paulraj, M., & Yaacob, S. (2008). Classification of Malaysian vowels using formant based features. Journal of Information and Communication Technology, 7, 27–40.

Zhou, X. J., & Dillon, T. S. A statistical-heuristic feature selection criterion for decision tree induction. 13 IEEE Transactions on Pattern Analysis and Machine Intelligence 834–841 (1991). IEEE Computer Society. doi:10.1109/34.85676

Downloads

Published

28-04-2015

How to Cite

Mohd Shaharanee, I. N., & Mohd Jamil, J. (2015). Irrelevant Feature and Rule Removal for Structural Associative Classification. Journal of Information and Communication Technology, 14, 95-110. https://doi.org/10.32890/jict2015.14.6

Research impact

Harvested 2026-09-06
4 citations, from OpenAlex — the highest of the sources checked

Counts differ between services because each indexes a different body of literature. None of them is the whole picture.

Identifiers DOI 10.32890/jict2015.14.6 OpenAlex W4254836248

Most read articles by the same author(s)