PolyGuard
Text dataset including annotated instructions or prompts for detecting biases, sensitive content, and model security.
Approximately 136,800 text examples annotated as lines in JSON/text format
CC-BY 4.0
Description
PolyGuard is a textual dataset composed of instructions and prompts annotated according to their biased or dangerous nature. Each example contains text and a label such as “unsafe” to report problematic content.
What is this dataset for?
- Train models for detecting biases and sensitive contents in NLP
- Improving the automated moderation of AI-generated content
- Contributing to ethics and safety in conversational AI systems
Can it be enriched or improved?
It is possible to enrich this dataset by adding new types of biases or sensitive content, as well as by refining the annotations for specific subcategories. Linguistic and cultural diversity can also be increased for greater robustness.
🔎 In summary
🧠 Recommended for
- NLP researchers
- Ethical AI teams
- Moderation developers
🔧 Compatible tools
- Hugging Face Transformers
- SpacY
- FastText
💡 Tip
Use data augmentation techniques to improve the robustness of the model in the face of subtle biases.
Frequently Asked Questions
What type of content does this dataset help detect?
It targets the detection of biased, dangerous, or inappropriate content in text prompts.
What license does this dataset cover?
It is licensed under CC-BY 4.0, allowing commercial and free use.
Is this dataset suitable for NLP beginners?
Yes, its simple structure and clear annotations make it easy for novices to use.




