Journal of Computational and Cognitive Engineering

← Back to Issue
Manuscript Framework

Enhanced Multimodal Webpage Classification Using Deep Learning for Efficient Information Retrieval

Volume
Volume 5
Issue Identifier
Issue No. 02
Publication Date
08 Sep 2025
export Digital Object Identifier

Abstract Scope

Web data mining has become a crucial tool for efficiently retrieving valuable information, as users increasingly rely on the World Wide Web for data exchange. Traditional web classification methods often struggle with handling multimodal data, leading to challenges in accurately classifying diverse web contents. Online classification plays a key role in facilitating efficient retrieval of information from multimedia content. This study presents a novel multimodal approach for webpage classification by integrating deep learning techniques for audio-visual analysis. The personalized Long Short-Term Memory (LS) TM model, which is a specific version of Long Short-Term Memory (LSTM), has improved classification accuracy by combining deep audio and video features. Artificial Convolutional Neural Networks (A-CNNs) extract complex audio features, while transformer networks capture long-range dependencies from video data. The present study proposes a log-sigmoid activation function that provides a more flexible thresholding method in logistic regression, thus greatly improving the classification performance. The focus of this study, which is on single-modality classification, presents an innovative method of integrating deep learning-based multimodal fusion, thus setting a new standard for web classification. Experimental results show that the Logistic Sigmoid Long Short-Term Memory ((LS)²TM) model has an accuracy of 88.09%, sensitivity of 89.14%, and specificity of 89.01%, outperforming state-of-the-art techniques such as LSTM, Deep Belief Network DBN, and A-CNN. The model also enhanced its precision (93.15%), recall (92.84%), and F-measure (93.25%), which are generally 5% higher than classical control methods. These findings highlight the potential of (LS)²TM for improving web content mining through multimodal analysis. Future research should focus on the real-world validation and scalability of dynamic web environments.