Intelligent Phishing Website Detection System Using Hybrid Machine Learning and Deep Learning Approaches

Uncategorized

Authors: Rajesh Chauhan, Akshay Bhardwaj, Shubham Rana

Abstract: Phishing websites are still one of the most common infection vectors for stealing credentials, financial information and other sensitive data from users of digital platforms, and there is a growing demand for accurate and deployable automated phishing-detection systems. However, models trained on a set of phishing URLs may experience a large performance drop when transferred to another feature scheme or a different collection window outside the scope of their training. In this work, we compare four classical machine-learning models, and five deep-learning and hybrid architectures for within-domain and cross-domain phishing-website detection. We evaluate our approach over five publicly available benchmark datasets, which contain more than 445k labeled instances in total, including UCI Phishing Websites Dataset, PhiUSIIL Phishing URL Dataset, Web-Page Phishing Detection Dataset, Mendeley Web-Page Phishing Dataset and Phishing Dataset based on Machine Learning. The models are evaluated in terms of classification accuracy, F1-score, transfer of cross domain performance, computational cost, data needs and interpretability. The Hybrid CNN-BiLSTM with attention model achieved the highest accuracy of 98.1% on PhiUSIIL and 85.6% on the transfer task from the PhiUSIIL to the Mendeley Web-Page dataset, outperforming the Linear SVM model by 13.1 percentage points. However, classical models still work well in resource-limited and interpretability-sensitive settings, and a simple BiLSTM with additive attention offers a balanced compromise. The results suggest that model selection for phishing detection should not only focus on predictive performance, but also consider operational deployment constraints.

DOI:

× How can I help you?