BullshiTeval — Assessing the reliability of AI assistants
Benchmark of 2,400 scenarios spread over 100 AI assistant models, used to detect and measure the “bullshit machine”.
Description
BullshitEval is a benchmark composed of 2,400 scenarios covering 100 AI assistants. It makes it possible to assess the quality of the answers provided by these models through various types of queries: general questions, flattery tests, negative concerns, etc. The dataset includes the context, the system prompt, the question, and the type of test.
What is this dataset for?
- Measuring the ability of LLMs to avoid inconsistent or erroneous responses
- Evaluate the behavior of AIs in the face of different types of prompts
- Building robust benchmarks for AI reliability research
Can it be enriched or improved?
Yes, BullshitEval can be adapted to other languages or expanded with new types of tests (e.g. hallucinations, cultural biases). It is also possible to add human annotations on the quality of responses or to include other AI models in the assessment.
🔎 In summary
🧠 Recommended for
- AI researchers
- LLM engineers
- Benchmark creators
🔧 Compatible tools
- Hugging Face Datasets
- Pandas
- Jupyter
- LLM Assessment Frameworks
💡 Tip
Use it with multiple LLMs in parallel to compare performance under the same conditions.
Frequently Asked Questions
Can we use BullshitEval to evaluate open-source models like Mistral or LLama?
Yes, the format is universal. You can test any model using the same prompts provided in the benchmark.
Does the dataset include high-quality human ratings?
No, model responses are not graded. Rather, it serves as a standardized basis for conducting your own evaluations.
Is the dataset suitable for testing French-language AI assistants?
The content is in English, but the scenarios can be translated and adapted for French-speaking AI assistants.




