By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
BullshiTeval — Assessing the reliability of AI assistants
Text

BullshiTeval — Assessing the reliability of AI assistants

Benchmark of 2,400 scenarios spread over 100 AI assistant models, used to detect and measure the “bullshit machine”.

Download dataset
Size

2,400 text examples in JSON format

Licence

MIT

Description

BullshitEval is a benchmark composed of 2,400 scenarios covering 100 AI assistants. It makes it possible to assess the quality of the answers provided by these models through various types of queries: general questions, flattery tests, negative concerns, etc. The dataset includes the context, the system prompt, the question, and the type of test.

What is this dataset for?

  • Measuring the ability of LLMs to avoid inconsistent or erroneous responses
  • Evaluate the behavior of AIs in the face of different types of prompts
  • Building robust benchmarks for AI reliability research

Can it be enriched or improved?

Yes, BullshitEval can be adapted to other languages or expanded with new types of tests (e.g. hallucinations, cultural biases). It is also possible to add human annotations on the quality of responses or to include other AI models in the assessment.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐⭐ (Very easy via Hugging Face Datasets)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (No cleaning required – ready-to-use data)
🏷️ Annotation richness⭐⭐⭐⭐✩ (Annotated by question type and system prompt type)
📜 Commercial license✅ Yes (MIT)
👨‍💻 Beginner friendly🌟 Yes, simple to explore and use
🔁 Fine-tuning ready⚠️ Not designed for fine-tuning, but useful for evaluation
🌍 Cultural diversity⚠️ Needs enrichment – currently focused on English

🧠 Recommended for

  • AI researchers
  • LLM engineers
  • Benchmark creators

🔧 Compatible tools

  • Hugging Face Datasets
  • Pandas
  • Jupyter
  • LLM Assessment Frameworks

💡 Tip

Use it with multiple LLMs in parallel to compare performance under the same conditions.

Frequently Asked Questions

Can we use BullshitEval to evaluate open-source models like Mistral or LLama?

Yes, the format is universal. You can test any model using the same prompts provided in the benchmark.

Does the dataset include high-quality human ratings?

No, model responses are not graded. Rather, it serves as a standardized basis for conducting your own evaluations.

Is the dataset suitable for testing French-language AI assistants?

The content is in English, but the scenarios can be translated and adapted for French-speaking AI assistants.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.