PhishSim: Aiding Phishing Website Detection with a Feature-Free Tool

Rizka Purwanto; Arindam Pal; Alan Blair; Sanjay Jha

arXiv:2207.10801·cs.CR·July 25, 2022

PhishSim: Aiding Phishing Website Detection with a Feature-Free Tool

Rizka Purwanto, Arindam Pal, Alan Blair, Sanjay Jha

PDF

TL;DR

This paper introduces PhishSim, a feature-free, compression-based method for detecting phishing websites that outperforms previous techniques with high accuracy and low false positives, suitable for real-time deployment.

Contribution

The paper presents a novel feature-free approach using Normalized Compression Distance and prototype selection for adaptive, efficient phishing detection without feature extraction.

Findings

01

Achieved 98.68% AUC in phishing detection

02

High TPR of around 90% with 0.58% FPR

03

Processing time of approximately 0.3 seconds

Abstract

In this paper, we propose a feature-free method for detecting phishing websites using the Normalized Compression Distance (NCD), a parameter-free similarity measure which computes the similarity of two websites by compressing them, thus eliminating the need to perform any feature extraction. It also removes any dependence on a specific set of website features. This method examines the HTML of webpages and computes their similarity with known phishing websites, in order to classify them. We use the Furthest Point First algorithm to perform phishing prototype extractions, in order to select instances that are representative of a cluster of phishing webpages. We also introduce the use of an incremental learning algorithm as a framework for continuous and adaptive detection without extracting new features when concept drift occurs. On a large dataset, our proposed method significantly…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.