What if a computer could tell a fake website is fake just by looking at its URL β without ever opening it? Turns out, yes. Here's how I built my first real machine learning project.
If you've ever gotten a message saying "Your bank account is suspended, click here now," you've seen phishing in action β fake websites disguised as real ones, built to steal your password or credit card number. Traditional defenses (blacklists of known bad URLs) are always playing catch-up. Attackers spin up new fake domains faster than blacklists can be updated.
Instead of memorizing bad URLs, what if a model learned the shared traits of phishing sites? Things like:
I used the UCI Phishing Websites dataset β 11,055 real websites, each labeled phishing or legitimate, described by 30 features.
I trained and compared three classic ML algorithms:
| Algorithm | Accuracy |
|---|---|
| Logistic Regression | 92.45% |
| Random Forest | 96.70% π |
| SVM | 94.71% |
When I checked which features mattered most, SSL certificate state and anchor link behavior dominated β by a wide margin over the other 28 features.
Why? Because phishing sites usually:
The model figured this out on its own β I never told it to focus on SSL.
The biggest lesson: you don't need to be an expert to start. The dataset was ready-made, the tools (Python + scikit-learn) are free, and the steps are well-documented. What it actually took was patience and consistency.
Full code and details are on GitHub:
π [https://github.com/eln2mac-has/phishing-detection-ml](https://github.com/eln2mac-has/phishing-detection-ml)
If you try something similar or have questions, drop a comment below π