Hashtag and Highest Scored Terms for Expanding Query

Authors

  • Wahyu Catur Wibowo Faculty of Computer Science University of Indonesia, Indonesia
  • Widodo Widodo Study Program of Informatics Engineering Education Universitas Negeri Jakarta, Indonesia

DOI:

https://doi.org/10.32890/jict2017.16.1.7

Keywords:

Query expansion, maximum hashtag, maximum term

Abstract

Communicating in short messages, such as using microblogs, was becoming more popular currently. Twitter https://twitter.com supports microblogs and retrieval of the blogs by users. To retrieve Twitter documents, we need specific strategies due to its specific characteristics. One new strategy for improving the effectiveness of twitter document retrieval is using the query expansion technique. This paper elaborates query expansion in twitter document retrieval by using the hashtag. We compared the effectiveness of query expansion in four different scenarios: the baseline result using no query expansion, highest scoredterm in terms of frequency-inverse document frequency (tfidf), maximum hashtag occurance, and combination of the highest scored-term and the maximum hashtag. The results show that the combination of the maximum term in tfidf and the maximum hashtag performs better in retrieving relevant documents than the baseline.

 

References

Aggarwal, N. & Buitelaar, P.(2012). Query expansion using Wikipedia and DBpedia. Retrived from: http://dblp.uni-trier.de/db/conf/clef/clef2012w

Damak, F., Pinel-Sauvagnat, K., & Cabanac, G.(2013). Effectiveness of State-of-the-art Features for Microblog Search”. SAC’13 March 18-22, 2013, Coimbra, Portugal.

Duan, Y, Jiang, L., Qin, T., Zhou, M., & Shum, HY.(2010). An empirical study on learning to rank tweets. In Proceedings of the 23rd International Conference on Computational Linguistics (Coling 2010), 295–303.

Efron, M. (2010). Hashtag retrieval in a microblogging environment. Proceedings of the 33rd International ACM SIGIR Conference on Research and Development in Information Retrieved,77-78. Switzerland: Geneva.

Hu, X., & Liu, H. (2012). Text analytics in social media. In Aggarwal, C. C., & Zhai, C. X. (Eds.), Mining Text Data. doi: 10.1007/978-1-4614-3223-4. Springer. Journal of ICT, 16, No. 1 (June) 2017, pp: 121-

Keskustalo, H., Kettunen, K., Kumpulainen, S., Ferro, N., Silvello, G., Jarvelin, A., Kekalainen, J., Arvola, P., Saastamoinen, M., & Jarvelin, K. (2015). Targeted query expansions as a method for searching: mixed quality digitized cultural heritage documents. iConference.

Lan, M., Tan, C.L., & Low, H.B. (2006). Proposing a new term weighting scheme for text categorization. Proceedings of the Twenty-First National Conference on Artifical Intelligence (AAAI-06), Boston, Massachusetts. pp. 763–768.

Leavitt, A., Burchard, E., Fisher, D., & Gilbert, S. (2009). The influentials: New approaches for analyzing influence on Twitter. Retrieved from http://www.webecologicalproject.org

Luo, Z., Osborne, M., Petrovic, S., & Wang, T. (2012). Improving Twitter retrieval by exploiting structural information. Proceedings of the Twenty-Sixth AAAI Conference on Artificial Intelligence.

Luo, Z., Yu, Y., Osborn, M., & Wang, T. (2015). Structuring Tweets for improving Twitter search. Journal of The Association for Information Science and Technology (ASIS&T).

Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval. USA: Cambridge University Press.

McCreadie, R. & Macdonald, C. (2013). Relevance in microblogs: Enhancing Tweet retrieval using hyperlinked documents. OAIR’2013, 22nd May, 2013, Lisbon, Portugal.

Nagmoti, R., Teredesai, A., & De Cock, M. (2010). Ranking approaches for microblog search. International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT).

Otsuka, E., Wallace, S.A., & Chiu, D. (2014). Design and evaluation of a twitter hashtag recommendation system. Proceedings of the 18th International Database Engineering & Applications Symposium, 330-333. Portugal: Porto.

Qiang, R., Fan, F., Lv, C., & Yang, J. (2015). Knowledge-based query expansion in real-time microblog search. The Asia Information Retrieval Societies Conference (AIRS). Journal of ICT, 16, No. 1 (June) 2017, pp: 121-

Shi, C., Xu, B., Lin, H., & Guo, Q. (2013). A time-sensitive model for microblog retrieval. NLPCC 2013, CCIS 400, pp. 402–409.

Tokunaga, T., & Iwayama, M. (1994). Text categorization based on weighted inverse document frequency. Techical Report 94-TR0001 Department of Computer Science Tokyo Institue of Technology.

Weerkamp, W., Balog, K., & De Rijke, M. (2012). Exploiting external collections for query expansion. ACM Transactions on the Web, 6(4), Article 18. doi: 10.1145/2382616.2.382621

Widodo & Wibowo, W.C. (2014).Improving classification performance by extending documents term. Proceedings of the International Conference on Data and Software Engineering (ICoDSE). ITB, Bandung.

Xuan, N. P., & Quang, H.L. (2014). A new improved term weighting scheme for text categorization. Knowledge and Systems Engineering 1, Advances In Intelligent Systems and Computing, 261-270. Springer.

Downloads

Published

31-05-2017

How to Cite

Wibowo, W. C., & Widodo, W. (2017). Hashtag and Highest Scored Terms for Expanding Query. Journal of Information and Communication Technology, 16(1), 121-135. https://doi.org/10.32890/jict2017.16.1.7

Research impact

Harvested 2026-09-06
0 citations recorded so far

Counts differ between services because each indexes a different body of literature. None of them is the whole picture.

Identifiers DOI 10.32890/jict2017.16.1.7 OpenAlex W4214767541