By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
Phishing URLs
Text

Phishing URLs

A balanced set of legitimate and fraudulent URLs, designed for phishing detection using automatic analysis methods.

Download dataset
Size

96,018 URLs in CSV format

Licence

Open Database License (ODbL)

Description

The dataset Phishing URLs contains 96,018 entries split evenly between legitimate URLs and phishing URLs. Each line is represented by a domain, a label (0 for legitimate, 1 for phishing) and a set of characteristics calculated according to metrics from the research article “PhishStorm”.

What is this dataset for?

  • Train classification models to detect fraudulent URLs
  • Testing web security algorithms in real time
  • Evaluate machine learning approaches on balanced binary data

Can it be enriched or improved?

Yes, it is possible to complete this data set with additional characteristics (e.g. WHOIS, IP geolocation, DNS history). Feature engineering techniques or temporal analysis models can also enrich the exploitation of the dataset.

🔎 In summary

Criterion Evaluation
🧩Ease of Use ⭐⭐⭐⭐⭐ (Easy to use, clean CSV format)
🧼Need for Cleaning ⭐⭐⭐☆☆ (Low – check for duplicates or anomalies)
🏷️Annotation Richness ⭐⭐⭐⭐☆ (Binary labels and multiple computed features)
📜Commercial License ✅ Yes (ODbL)
👨‍💻Beginner Friendly 🧑‍💻 Yes – good foundation for ML cybersecurity projects
🔁Reusable for Fine-tuning 🛠️ Usable for pre-training or benchmarking
🌍Cultural Diversity 🌍 Not specified but potentially global

🧠 Recommended for

  • Cybersecurity projects
  • Phishing detection training
  • Binary classification

🔧 Compatible tools

  • Scikit-learn
  • Pandas
  • LightGBM
  • XGBoost
  • TensorFlow

💡 Tip

Use balanced distribution to experiment with different evaluation metrics like AUC, F1-score, and precision/recall.

Frequently Asked Questions

Can this dataset be used to train a model in production?

Yes, as long as you enrich it with more recent data and adapt the features to the context of your web service.

Is the dataset balanced between phishing and non-phishing?

Yes, it contains exactly 48,009 legitimate URLs and 48,009 phishing URLs.

Is it suitable for supervised learning?

In fact, the dataset is labeled and formatted for training classic or advanced supervised models.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.