By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
StreamUni
Audio

StreamUni

Audio dataset used to train a continuous speech translation model. Great for speech translation streaming tasks.

Download dataset
Size

9,631 audio files, Parquet format (~693 MB)

Licence

Apache 2.0

Description

‍

StreamUni is a structured audio dataset for continuous voice translation tasks. It contains 9,631 audio samples with associated transcripts and translations, used to train the StreamUni-phi4 model. The dataset is compact (~693 MB) and provides a solid basis for streaming speech translation searches.

‍

‍

What is this dataset for?

‍

  • Train models capable of translating speech live (streaming speech translation)
  • Testing the effectiveness of multilingual transcription/translation systems in real time
  • Serve as a base for voice assistants or automatic subtitling applications

‍

‍

Can it be enriched or improved?

‍

Yes. It is possible to increase this dataset with additional languages, audio metadata (duration, accent, speaker) or even annotations on fluidity/translatability. Finer segmentation could also improve performance in streaming mode.

‍

‍

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐✩ (Easy to use with Hugging Face)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low – well-prepared data)
🏷️ Annotation richness⭐⭐⭐✩✩ (Transcriptions and translations available)
📜 Commercial license✅ Yes (Apache 2.0)
👨‍💻 Beginner friendly⚠️ Requires some audio/linguistic knowledge
🔁 Fine-tuning ready🎯 Highly suitable for speech-to-text model fine-tuning
🌍 Cultural diversity⚠️ Medium – could be enriched with more languages or accents

‍

‍

🧠 Recommended for

  • NLP speech engineers
  • Voice translation projects
  • Automatic subtitling prototyping

‍

‍

🔧 Compatible tools

  • Hugging Face Datasets
  • Whisper
  • ESPnet
  • Fairseq

‍

‍

💡 Tip

For smoother translation, cut long files into shorter speech units.

Frequently Asked Questions

Is this dataset suitable for offline translation?

It was designed for streaming, but can also be used in offline audio translation scenarios.

Can it be easily integrated into a Hugging Face pipeline?

Yes, it is available in Parquet format and compatible with the Hugging Face `datasets` library.

Is it multilingual?

It includes multilingual data, but the set of languages covered depends on the model used for the initial generation.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.