A New Approach for Video Concept Detection Based on User Comments
DOI:
https://doi.org/10.32890/jict2021.20.4.7Keywords:
based video retrieval, social media tagging, natural language processing, video concept detectionAbstract
Video concept detection means describing a video with semantic concepts that correspond to the content of the video. The concepts
help to retrieve video quickly. These semantic concepts describe high-level elements that depict the key information present in the
content. In recent years, many efforts have been done to automate this task because the manual solution is time-consuming. Nowadays, videos come with comments. Therefore, in addition to the content of the videos, the comments should be analyzed because they contain valuable data that help to retrieve videos. This paper focused especially on videos shared on social media. The specificity of these videos was the presence of massive comments. This paper attempted to exploit comments by extracting concepts from them. This would support the research effort that works only on the visual content. Natural language processing techniques were used to analyze comments and to filter words to retain only the ones that could be considered as concepts. The proposed approach was tested on YouTube videos. The results demonstrated that the proposed approach was able to extract accurate data and concepts from the comments that could be used to ease the retrieval of videos. The findings supported the research effort of working on the visual and audio contents of the videos.
References
Anthony, L. (2021). Computer Software. Waseda University. https://www.laurenceanthony.net/software/tagant/
Awad, G., Fiscus, J., Joy, D., Michel, M., Smeaton, A., Kraaij, W.,... & Larson, M. (2016, November). Trecvid 2016: Evaluating video search, video event detection, localization, and hyperlinking. In TREC Video Retrieval Evaluation (TRECVID).
Ballan, L., Bertini, M., Uricchio, T., & Del Bimbo, A. (2015). Data-driven approaches for social image and video tagging. Multimedia Tools and Applications, 74(4), 1443–1468.
Cagliero, L., Canale, L., & Farinetti, L. (2019, July). VISA: A supervised approach to indexing video lectures with semantic annotations. In 2019 IEEE 43rd Annual Computer Software and Applications Conference (COMPSAC) (Vol. 1, pp. 226–235). IEEE. Google I/O 2013 - Semantic video annotations in the Youtube topics API: Theory and applications [online]. Available: https://www. youtube.com/watch?v=wf_77z1H-vQ
Huurnink, B., Snoek, C. G., de Rijke, M., & Smeulders, A. W. (2012). Content-based analysis improves audiovisual archive retrieval. IEEE Transactions on Multimedia, 14(4), 1166–1178. Journal of ICT, 20, No. 4 (October) 2021, pp: 629–
Klami, M., & Lagus, K. (2006, September). Unsupervised word categorization using self-organizing maps and automatically extracted morphs. In International Conference on Intelligent Data Engineering and Automated Learning (pp. 912–919). Springer.
Kohonen, T. (1982). Self-organized formation of topologically correct feature maps. Biological Cybernetics, 43(1), 59–69.
Kordumova, S., Li, X., & Snoek, C. G. (2015). Best practices for learning video concept detectors from social media examples. Multimedia Tools and Applications, 74(4), 1291–1315.
Landi, F., Snoek, C. G., & Cucchiara, R. (2019). Anomaly locality in video surveillance. arXiv preprint arXiv:1901.10364.
Li, H., Shi, Y., Liu, Y., Hauptmann, A. G., & Xiong, Z. (2012). Cross-domain video concept detection: A joint discriminative and generative active learning approach. Expert Systems with Applications, 39(15), 12220–12228.
Li, X., Snoek, C. G., & Worring, M. (2008, October). Learning tag relevance by neighbor voting for social image retrieval. In Proceedings of the 1st ACM International Conference on Multimedia Information Retrieval (pp. 180–187).
Liu, S., & Forss, T. (2015, November). Automatic tag extraction from social media for visual labeling. In 2015 7th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management (IC3K) (Vol. 1, pp. 504–510). IEEE.
Mansouri, S., Charhad, M., Rekik, A., & Zrigui, M. (2018, August). A framework for semantic video content indexing using textual information. In 2018 IEEE Second International Conference on Data Stream Mining & Processing (DSMP) (pp. 107–110). IEEE.
McGuinness, K., Mohedano, E., Salvador, A., Zhang, Z., Marsden, M., Wang, P., & Smeaton, A., (2015). Insight Dublin City University at Text Retrieval Conference-Video Track 2015. In Text Retrieval Conference - Video Track 2015 Overview Papers and Slidespp (pp. 1–16).
Miller, G. (1995). WordNet: A lexical database for English. Communications of the Association for Computing Machinery, 38(11), 39–41.
Mohedano, E., McGuinness, K., O’Connor, N. E., Salvador, A., Marques, F., & Giró-i-Nieto, X. (2016, June). Bags of local convolutional features for scalable instance search. Journal of ICT, 20, No. 4 (October) 2021, pp: 629– In Proceedings of the 2016 ACM on International Conference on Multimedia Retrieval (327–331).
Mettes, P., Snoek, C. G., & Chang, S. F. (2017). Localizing actions from video labels and pseudo-annotations. arXiv preprint arXiv:1707.09143.
Muchtar, N., Afdhal K., Dwiyantoro A., & Prayuda A. (2020). Deep anomaly detection through visual attention in surveillance videos. Journal of Big Data, 7(87).
Pacella, M., Grieco, A., & Blaco, M. (2016). On the use of self-organizing map for text clustering in engineering change process analysis: A case study. Computational Intelligence And Neuroscience, 2016.
Peixoto, B., Lavi, B., Bestagini, P., Dias, Z., & Rocha, A. (2020, May). Multimodal violence detection in videos. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 2957–2961). IEEE.
Safadi, B., Derbas, N., Hamadi, A., Budnik, M., Mulhem, P., & Quénot, G. (2014, November). LIG at TRECVid 2014: Semantic indexing. In Proceedings of TRECVID.
Safadi, B., Derbas, N., & Quénot, G. (2015). Descriptor optimization for multimedia indexing and retrieval. Multimedia Tools and Applications, 74(4), 1267–1290.
Sawant, N., Datta, R., Li, J., & Wang, J. Z. (2010, March). Quest for relevant tags using local interaction networks and visual content. In Proceedings of the International Conference on Multimedia Information Retrieval (pp. 231–240).
Salton G., & McGill, M. J. (1986). Introduction to Modern Information Retrieval. McGraw-Hill.
Snoek, C. G., Cappallo, S., Fontijne, D., Julian, D., Koelma, D. C., Mettes, P.,... & Towal, R. B. (2015). Qualcomm Research and University of Amsterdam at TRECVID 2015: Recognizing concepts, objects, and events in video. In TRECVID.
Snoek, C. G., Li, X., Xu, C., & Koelma, D. C. (2017). University of Amsterdam and Renmin University at TRECVID 2017: Searching video, detecting events and describing video. In TRECVID.
Snoek, C. G., Worring M., Geusebroek J., Koelma D. C., Seinstra F. J., & Smeulders A. W. (2008). The semantic pathfinder: Using an authoring metaphor for generic multimedia indexing. IEEE Transactions on Pattern Analysis and Machine Intelligence, 28(2006), 1678–1689. Journal of ICT, 20, No. 4 (October) 2021, pp: 629–
Resnik, P. (1999). Semantic similarity in a taxonomy: An information based measure and its application to problems of ambiguity in natural language. Journal of Artificial Intelligence Research, 11, 95–130.
Ultsch, A., & Siemon, H. P. (1990, July). Kohonen’s self-organizing feature maps for exploratory data analysis. In Proceedings of the International Neural Network Conference. Kluwer Academic Press.
Uricchio, T., Ballan, L., Seidenari, L., & Del Bimbo, A. (2017). Automatic image annotation via label transfer in the semantic space. Pattern Recognition, 71, 144–157.
Wu, Z., & Palmer, M. (1994). Verbs semantics and lexical selection. In Proceedings of the 32nd annual meeting on Association for Computational Linguistics (pp. 133–138). YouTube Academy-Last Visited, April 2020: https://creatoracademy. youtube.com/page/home?hl=fr
Yu, S.I., Jiang, L., Xu, Z., Lan, Z., Xu, S., Chang, X., Li, X., Mao, Z., Gan, C., & Miao, Y. (2015). Informedia@ text retrieval conference - Video track 2015 MED. In National Institute of Standards and Technology, Text Retrieval Conference - Video Track Workshop.
Zhu, G., Yan, S., & Ma, Y. (2010, October). Image tag refinement towards low-rank, content-tag prior and error sparsity. In Proceedings of the 18th ACM international conference on Multimedia (pp. 461–470).
Published
Issue
Section
License
Copyright (c) 2022 Journal of Information and Communication Technology

This work is licensed under a Creative Commons Attribution 4.0 International License.
How to Cite
Research impact
Harvested 2026-09-06Counts differ between services because each indexes a different body of literature. None of them is the whole picture.
2002 - 2020






















