Phishing Email Dataset
Dataset compiling more than 82,000 emails from multiple sources, annotated for phishing detection. Emails contain subject, message body, and associated metadata. This corpus is suitable for training models for the classification and detection of phishing attacks.
Approximately 82,500 emails in text format, with metadata (sender, recipient, date), annotated spam/phishing or legitimate
CC BY-SA 4.0
Description
The dataset Phishing Email Dataset includes approximately 82,500 emails annotated as spam/phishing or legitimate, from various databases (Enron, Ling, SpamAssassin, etc.). It contains the subject, body of the message, and metadata such as sender, recipient, and date.
What is this dataset for?
- Train automatic phishing and spam detection models through text analysis.
- Study the tactics and characteristics of fraudulent emails.
- Develop filtering and security tools for email.
Can it be enriched or improved?
Yes, it is possible to add additional annotations, such as classification by type of phishing attack or the detection of specific malicious elements. The dataset can also be updated regularly with new campaigns.
🔎 In summary
🧠 Recommended for
- Cybersecurity researchers
- Spam filter developers
- NLP students
🔧 Compatible tools
- TensorFlow
- PyTorch
- SpacY
- Scikit-learn
- Email parsing libraries
💡 Tip
Test hybrid approaches combining textual analysis and metadata to improve detection.
Frequently Asked Questions
Does this dataset contain recent emails that are representative of current phishing tactics?
The dataset combines several traditional sources, it is advisable to complete with more recent data to stay up to date.
Do emails contain attachments or only text?
Mostly text with subject and body, attachments are not included.
Can I use this dataset for a commercial project?
Yes, the CC BY-SA 4.0 license allows commercial use under identical sharing conditions.



