A Deep Autoencoder-Based Representation for Arabic Text Categorization
DOI:
https://doi.org/10.32890/jict2020.19.3.10Keywords:
Arabic text representation, deep autoencoder, feature selection, machine learning, text categorizationAbstract
Arabic text representation is a challenging assignment for several applications such as text categorization and clustering since the Arabic language is known for its variety, richness and complex morphology. Until recently, the Bag-of-Words remains the most common method for Arabic text representation. However, it suffers from several shortcomings such as semantics deficiency and high dimensionality of feature space. Moreover, most existing methods ignore the explicit knowledge contained in semantic vocabularies such as Arabic WordNet. To overcome these shortcomings, we proposed a deep Autoencoder based representation for Arabic text categorization. It consisted of three stages: (1) Extracting from Arabic WordNet the most relevant concepts based on feature selection processes (2) Features learning via an unsupervised algorithm for text representation (3) Categorizing text using deep Autoencoder. Our method allowed for the consideration of document semantics by combining both implicit and explicit semantics and reducing feature space dimensionality. To evaluate our method, we conducted several experiments on the standard Arabic dataset, OSAC. The obtained results showed the effectiveness of the proposed method compared to state-of-the-art ones.
References
Abdullah, M., & Shaikh, S. (2018, June). Teamuncc at semeval-2018 task 1: Emotion detection in English and Arabic tweets using deep learning. In Proceedings of the 12th International Workshop on Semantic Evaluation (pp. 350-357).
Abu-Errub, A. (2014). Arabic Text Classification Algorithm using TFIDF and Chi Square Measurements. International Journal of Computer Applications, 93(6).
Alayba, A. M., Palade, V., England, M., & Iqbal, R. (2018a, March). Improving sentiment analysis in Arabic using word representation. In 2018 IEEE 2nd International Workshop on Arabic and Derived Script Analysis and Recognition (ASAR) (pp. 13-18). IEEE.
Alayba, A. M., Palade, V., England, M., & Iqbal, R. (2018b, August). A combined CNN and LSTM model for Arabic sentiment analysis. In International Cross-Domain Conference for Machine Learning and Knowledge Extraction (pp. 179-191). Springer, Cham.
Al-Anzi, F. S., & AbuZeina, D. (2017). Toward an enhanced Arabic text classification using cosine similarity and Latent Semantic Indexing. Journal of King Saud University - Computer and Information Sciences, 29(2), 189-195.
Al-Salemi, B., Ayob, M., Kendall, G., & Noah, S. A. M. (2019). Multi-label Arabic text categorization: A benchmark and baseline comparison of multi-label learning algorithms. Information Processing & Management, 56(1), 212-227.
Al-Sallab, A., Baly, R., Hajj, H., Shaban, K. B., El-Hajj, W., & Badaro, G. (2017). Aroma: A recursive deep learning model for opinion mining in Arabic as a low resource language. ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP), 16(4), 1-20.
Al-Smadi, M., Qawasmeh, O., Al-Ayyoub, M., Jararweh, Y., & Gupta, B. (2018). Deep Recurrent neural network vs. support vector machine for aspect-based sentiment analysis of Arabic hotels’ reviews. Journal of Computational Science, 27, 386-393.
Al-Smadi, M., Talafha, B., Al-Ayyoub, M., & Jararweh, Y. (2019). Using long short-term memory deep neural networks for aspect-based sentiment analysis of Arabic reviews. International Journal of Machine Learning and Cybernetics, 10(8), 2163-2175.
Bengio, Y., Schwenk, H., Senécal, J.-S., Morin, F., & Gauvain, J.-L. (2006). Neural Probabilistic Language Models. Innovations in Machine Learning, 137-186.
Carreira-Perpinan, M. A., & Hinton, G. E. (2005). On Contrastive Divergence Learning. In Aistats (Vol. 10, pp. 33-40). Journal of ICT, 19, No. 3 (July) 2020, pp: 381-
El Mahdaouy, A., Gaussier, E., & El Alaoui, S. O. (2016, October). Arabic text classification based on word and document embeddings. In International Conference on Advanced Intelligent Systems and Informatics (pp. 3241). Springer, Cham.
Elnagar, A., Al-Debsi, R., & Einea, O. (2020). Arabic text classification using deep learning models. Information Processing & Management, 57(1), 102-121.
El-Alami, F. Z., & El Alaoui, S. O. (2016, December). An Efficient Method based on Deep Learning Approach for Arabic Text Categorization. In International Arab Conference on Information Technology.
El-Alami, F. Z., & El Alaoui, S. O. (2018, November). Word sense representation based-method for Arabic text categorization. In 2018 9th International Symposium on Signal, Image, Video and Communications (ISIVC) (pp. 141-146). IEEE.
Guyon, I., & Elisseeff, A. (2003). An introduction to variable and feature selection. Journal of Machine Learning Research, 3(Mar), 1157-1182.
Hinton, G. E., & Salakhutdinov, R. R. (2006). Reducing the Dimensionality of
Data with Neural Networks. Science, 313(5786), 504-507.
Le, Q., & Mikolov, T. (2014, January). Distributed representations of sentences and documents. In International Conference on Machine Learning (pp. 1188-1196).
Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013a). Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., & Dean, J. (2013b). Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems (pp. 3111-3119).
Odeh, A., Abu-Errub, A., Shambour, Q., & Turab, N. (2015). Arabic Text CategorizationAlgorithm Using Vector Evaluation Method. International Journal of Computer Science and Information Technology, 6(6), 8392.
Salakhutdinov, R., & Hinton, G. (2009). Semantic hashing. International Journal of Approximate Reasoning, 50(7), 969-978.
Swesi, I. M. A. O. & Bakar, A. B. (2019). Feature clustering for pso-based feature construction on high-dimensional data. Journal of Information and Communication Technology, 18(4), 439-472.
Tan, C. C., & Eswaran, C. (2008, May). Performance comparison of three types of autoencoder neural networks. In 2008 Second Asia International Conference on Modelling & Simulation (AMS) (pp. 213-218). IEEE.
Tian, X., Hérault, R., Gasso, G., & Canu, S. (2010, January). Pré-apprentissage supervisé pour les réseaux profonds. In Proceedings of Rfia (Vol. 2010, p. 36). Journal of ICT, 19, No. 3 (July) 2020, pp: 381-
Yang, Y., & Pedersen, J. O. (1997, July). A comparative study on feature selection in text categorization. International Conference on Machine Learning (Vol. 97, No. 412-420, p. 35).
Yousif, S. A., Samawi, V. W., Elkabani, I., & Zantout, R. (2015). The Effect of Combining Different Semantic Relations on Arabic Text Classification. World of Computer Science & Information Technology Journal, 5(1), 12-118.
Zrigui, M., Ayadi, R., Mars, M., & Maraoui, M. (2012). Arabic Text Classification Framework Based on Latent Dirichlet Allocation. Journal of Computing and Information Technology, 20(2), 125-140.
Published
Issue
Section
How to Cite
Research impact
Harvested 2026-09-25Counts differ between services because each indexes a different body of literature. None of them is the whole picture.
2002 - 2020






















