AudioTrust: Benchmark AllMS
AudioTrust is a large-scale benchmark designed to assess audio language models (AllMS) on six critical axes: hallucination, security, robustness, fairness, privacy, and resistance to attacks. It provides standardized tasks and metrics to test the true reliability of audio AI systems.
Audio format + structured annotations, multi-task benchmark (6 dimensions), WAV + JSON files
CC-BY-SA 4.0
Description
AudioTrust is an evaluation benchmark designed to thoroughly test multimodal audio models (AllMS). It covers six essential dimensions: detection of hallucinations, robustness to corruption, voice spoofing tests, private data leaks, demographic equity, and the generation of dangerous content.
What is this dataset for?
- Evaluate the resistance of an audio model to vocal cloning attacks
- Test audio and text generation under security constraints
- Audit demographic biases in generated voice responses
Can it be enriched or improved?
Yes, AudioTrust can be complemented by additional voices from different languages or demographics. It is also possible to add new attack scenarios, tasks in non-English languages, or to combine tests with human evaluations for more robust results.
🔎 In summary
🧠 Recommended for
- AI security researchers
- Voice authentication tool developers
- Audio model listeners
🔧 Compatible tools
- PyTorch
- Whisper
- Transformers
- DeepSpeech
- Librosa
💡 Tip
To finely test robustness, inject simulated audio disturbances with realistic noises or distortions before evaluation.
Frequently Asked Questions
Is this dataset intended for training audio models?
No, it is designed for evaluating existing models based on trust criteria, not for fine-tuning.
What types of models can we test with AudioTrust?
It is designed for multimodal audio models (AllMS), but can also be used to test ASR models, vocoders, or voice encoders.
Does the dataset contain real or synthetic voice data?
It combines real and scripted recordings to simulate attacks and various test conditions.




