Analysis of the Performance Impact of Fine-Tuned Machine Learning Model for Phishing URL Detection

Phishing leverages people’s tendency to share personal information online. Phishing attacks often begin with an email and can be used for a variety of purposes. The cybercriminal will employ social engineering techniques to get the target to click on the link in the phishing email, which will take t...

Full description

Saved in:
Bibliographic Details
Main Author: Abdul Samad, Saleem Raja (author)
Other Authors: Al-Kaabi, Amna Salim (author), Balasubaramanian, Sundarvadivazhagan (author), Bostani, Ali (author), Chowdhury, Subrata (author), Mehbodniya, Abolfazl (author), Sharma, Bhisham (author), Webber, Julian L. (author)
Published: 2023
Online Access:http://hdl.handle.net/11675/10946
http://www.scopus.com/inward/record.url?scp=85152791137&partnerID=8YFLogxK
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1870679722260692992
author Abdul Samad, Saleem Raja
author2 Al-Kaabi, Amna Salim
Balasubaramanian, Sundarvadivazhagan
Bostani, Ali
Chowdhury, Subrata
Mehbodniya, Abolfazl
Sharma, Bhisham
Webber, Julian L.
author2_role author
author
author
author
author
author
author
author_facet Abdul Samad, Saleem Raja
Al-Kaabi, Amna Salim
Balasubaramanian, Sundarvadivazhagan
Bostani, Ali
Chowdhury, Subrata
Mehbodniya, Abolfazl
Sharma, Bhisham
Webber, Julian L.
author_role author
dc.creator.none.fl_str_mv Abdul Samad, Saleem Raja
Al-Kaabi, Amna Salim
Balasubaramanian, Sundarvadivazhagan
Bostani, Ali
Chowdhury, Subrata
Mehbodniya, Abolfazl
Sharma, Bhisham
Webber, Julian L.
dc.date.none.fl_str_mv 2023-04-01
2024-02-05T08:32:42Z
2024-02-05T08:32:42Z
dc.identifier.none.fl_str_mv 10.3390/electronics12071642
http://hdl.handle.net/11675/10946
http://www.scopus.com/inward/record.url?scp=85152791137&partnerID=8YFLogxK
dc.publisher.none.fl_str_mv Multidisciplinary Digital Publishing Institute (MDPI)
dc.relation.none.fl_str_mv Electrical and Computer Engineering
Electronics (Switzerland)
dc.title.none.fl_str_mv Analysis of the Performance Impact of Fine-Tuned Machine Learning Model for Phishing URL Detection
dc.type.none.fl_str_mv Journal Article
info:eu-repo/semantics/publishedVersion
description Phishing leverages people’s tendency to share personal information online. Phishing attacks often begin with an email and can be used for a variety of purposes. The cybercriminal will employ social engineering techniques to get the target to click on the link in the phishing email, which will take them to the infected website. These attacks become more complex as hackers personalize their fraud and provide convincing messages. Phishing with a malicious URL is an advanced kind of cybercrime. It might be challenging even for cautious users to spot phishing URLs. The researchers displayed different techniques to address this challenge. Machine learning models improve detection by using URLs, web page content and external features. This article presents the findings of an experimental study that attempted to enhance the performance of machine learning models to obtain improved accuracy for the two phishing datasets that are used the most commonly. Three distinct types of tuning factors are utilized, including data balancing, hyper-parameter optimization and feature selection. The experiment utilizes the eight most prevalent machine learning methods and two distinct datasets obtained from online sources, such as the UCI repository and the Mendeley repository. The result demonstrates that data balance improves accuracy marginally, whereas hyperparameter adjustment and feature selection improve accuracy significantly. The performance of machine learning algorithms is improved by combining all fine-tuned factors, outperforming existing research works. The result shows that tuning factors enhance the efficiency of machine learning algorithms. For Dataset-1, Random Forest (RF) and Gradient Boosting (XGB) achieve accuracy rates of 97.44% and 97.47%, respectively. Gradient Boosting (GB) and Extreme Gradient Boosting (XGB) achieve accuracy values of 98.27% and 98.21%, respectively, for Dataset-2.
id AUKR_44dbc515f2175f8662fc0f4da7ac84bb
identifier_str_mv 10.3390/electronics12071642
network_acronym_str AUKR
network_name_str AU Kuwait Rep
oai_identifier_str oai:dspace.auk.edu.kw:11675/10946
publishDate 2023
publisher.none.fl_str_mv Multidisciplinary Digital Publishing Institute (MDPI)
repository.mail.fl_str_mv
repository.name.fl_str_mv
repository_id_str
spelling Analysis of the Performance Impact of Fine-Tuned Machine Learning Model for Phishing URL DetectionAbdul Samad, Saleem RajaAl-Kaabi, Amna SalimBalasubaramanian, SundarvadivazhaganBostani, AliChowdhury, SubrataMehbodniya, AbolfazlSharma, BhishamWebber, Julian L.Phishing leverages people’s tendency to share personal information online. Phishing attacks often begin with an email and can be used for a variety of purposes. The cybercriminal will employ social engineering techniques to get the target to click on the link in the phishing email, which will take them to the infected website. These attacks become more complex as hackers personalize their fraud and provide convincing messages. Phishing with a malicious URL is an advanced kind of cybercrime. It might be challenging even for cautious users to spot phishing URLs. The researchers displayed different techniques to address this challenge. Machine learning models improve detection by using URLs, web page content and external features. This article presents the findings of an experimental study that attempted to enhance the performance of machine learning models to obtain improved accuracy for the two phishing datasets that are used the most commonly. Three distinct types of tuning factors are utilized, including data balancing, hyper-parameter optimization and feature selection. The experiment utilizes the eight most prevalent machine learning methods and two distinct datasets obtained from online sources, such as the UCI repository and the Mendeley repository. The result demonstrates that data balance improves accuracy marginally, whereas hyperparameter adjustment and feature selection improve accuracy significantly. The performance of machine learning algorithms is improved by combining all fine-tuned factors, outperforming existing research works. The result shows that tuning factors enhance the efficiency of machine learning algorithms. For Dataset-1, Random Forest (RF) and Gradient Boosting (XGB) achieve accuracy rates of 97.44% and 97.47%, respectively. Gradient Boosting (GB) and Extreme Gradient Boosting (XGB) achieve accuracy values of 98.27% and 98.21%, respectively, for Dataset-2.Multidisciplinary Digital Publishing Institute (MDPI)2024-02-05T08:32:42Z2024-02-05T08:32:42Z2023-04-01Journal Articleinfo:eu-repo/semantics/publishedVersion10.3390/electronics12071642http://hdl.handle.net/11675/10946http://www.scopus.com/inward/record.url?scp=85152791137&partnerID=8YFLogxKElectrical and Computer EngineeringElectronics (Switzerland)oai:dspace.auk.edu.kw:11675/109462024-02-05T08:32:42Z
spellingShingle Analysis of the Performance Impact of Fine-Tuned Machine Learning Model for Phishing URL Detection
Abdul Samad, Saleem Raja
status_str publishedVersion
title Analysis of the Performance Impact of Fine-Tuned Machine Learning Model for Phishing URL Detection
title_full Analysis of the Performance Impact of Fine-Tuned Machine Learning Model for Phishing URL Detection
title_fullStr Analysis of the Performance Impact of Fine-Tuned Machine Learning Model for Phishing URL Detection
title_full_unstemmed Analysis of the Performance Impact of Fine-Tuned Machine Learning Model for Phishing URL Detection
title_short Analysis of the Performance Impact of Fine-Tuned Machine Learning Model for Phishing URL Detection
title_sort Analysis of the Performance Impact of Fine-Tuned Machine Learning Model for Phishing URL Detection
url http://hdl.handle.net/11675/10946
http://www.scopus.com/inward/record.url?scp=85152791137&partnerID=8YFLogxK