A Comparative Analysis of Feature Selection Methods for Machine Learning-Based Phishing URL Detection

Authors

  • Farah Yulianti Universitas President
  • Jessica Forenziana Universitas President

DOI:

https://doi.org/10.55927/fjcis.v5i2.16783

Keywords:

Phishing Detection, Feature Selection, Machine Learning, URL Classification, Cybersecurity

Abstract

This study investigates the impact of three categories of feature selection methods which are filter-based, wrapper-based, and embedded on the performance and efficiency of machine learning classifiers for phishing URL detection. Experiments were conducted on the PhiUSIIL Phishing URL Dataset comprising 235,795 URLs (134,850 legitimate, 100,945 phishing) with 20 URL structural features. Results demonstrate that LASSO (Embedded) at k = 5 achieves an F1 score of 0.9975 versus the 20-feature baseline of 0.9965 (ΔF1 = 0.001). Given that this difference falls within the per-fold standard deviation of both configurations (≤ 0.0003), we treat the two configurations as operationally equivalent and frame the contribution as efficiency at unchanged accuracy rather than as an accuracy improvement.

Downloads

Download data is not yet available.

References

Anti-Phishing Working Group. (2025). Phishing Activity Trends Report, 4th Quarter 2024. Published 19 March 2025. Retrieved from https://apwg.org/trendsreports.

Bari, N., Saleem, T., Shah, M., Algarni, A., Patel, A., & Ullah, I. (2025). A filter-based feature selection framework to detect phishing URLs using stacking ensemble machine learning. Computer Modeling in Engineering & Sciences, 145(1), 1167–1187.

Biggio, B., & Roli, F. (2018). Wild patterns: Ten years after the rise of adversarial machine learning. Pattern Recognition, 84, 317–331. https://doi.org/10.1016/j.patcog.2018.07.023.

Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32.

Calzarossa, M. C., Giudici, P., & Zieni, R. (2024). Explainable machine learning for phishing feature detection. Quality and Reliability Engineering International, 40(1), 362–373. https://doi.org/10.1002/qre.3411.

Capuano, N., Fenza, G., Loia, V., & Stanzione, C. (2022). Explainable artificial intelligence in cybersecurity: A survey. IEEE Access, 10, 93575–93600. https://doi.org/10.1109/ACCESS.2022.3204171.

Federal Bureau of Investigation, Internet Crime Complaint Center. (2025). 2024 Internet Crime Report. Retrieved from https://www.ic3.gov/AnnualReport/Reports/2024_IC3Report.pdf.

Fernando, W., & Komninos, N. (2026). Evaluating and combating the impact of concept drift on the performance of machine learning-based phishing detection systems. arXiv preprint arXiv:2606.11471.

Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. Annals of Statistics, 29(5), 1189–1232.

Guyon, I., Weston, J., Barnhill, S., & Vapnik, V. (2002). Gene selection for cancer classification using support vector machines. Machine Learning, 46(1–3), 389–422.

Khonji, M., Iraqi, Y., & Jones, A. (2013). Phishing detection: A literature survey. IEEE Communications Surveys & Tutorials, 15(4), 2091–2121.

Kocyigit, E., Korkmaz, M., Sahingoz, O. K., & Diri, B. (2024). Enhanced feature selection using genetic algorithm for machine-learning-based phishing URL detection. Applied Sciences, 14(14), 6081. https://doi.org/10.3390/app14146081.

Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30, 4765–4774.

Pal, S., Yadav, G., Jadidi, Z., Habib, A., Uddin, M. P., Karmakar, C., & Shukla, S. (2025). Vulnerabilities in machine learning for cybersecurity: Current trends and future research directions. Journal of Information Security and Applications, 96, 104269.

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., ... & Duchesnay, E. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.

Prasad, A., & Chandra, S. (2024). PhiUSIIL: A diverse security profile empowered phishing URL detection framework based on similarity index and incremental learning. Computers & Security, 136, 103545.

Sahingoz, O. K., Buber, E., Demir, O., & Diri, B. (2019). Machine learning based phishing detection from URLs. Expert Systems with Applications, 117, 345–357. https://doi.org/10.1016/j.eswa.2018.09.029.

Sarker, I. H., Kayes, A. S. M., Badsha, S., Alqahtani, H., Watters, P., & Ng, A. (2020). Cybersecurity data science: An overview from machine learning perspective. Journal of Big Data, 7(1), 1–25.

Tibshirani, R. (1996). Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society: Series B, 58(1), 267–288.

Uddin, K. M., Biswas, N., Rikta, S. T., et al. (2025). Explainable machine learning for phishing site detection: A high-efficiency approach using boosting models and SHAP. The Journal of Engineering, 2025(7), e70110. https://doi.org/10.1049/tje2.70110.

Published

2026-08-26

How to Cite

Yulianti, F., & Forenziana, J. (2026). A Comparative Analysis of Feature Selection Methods for Machine Learning-Based Phishing URL Detection. Formosa Journal of Computer and Information Science, 5(2), 243–258. https://doi.org/10.55927/fjcis.v5i2.16783

Issue

Section

Articles