StreamUni
Audio dataset used to train a continuous speech translation model. Great for speech translation streaming tasks.
Description
StreamUni is a structured audio dataset for continuous voice translation tasks. It contains 9,631 audio samples with associated transcripts and translations, used to train the StreamUni-phi4 model. The dataset is compact (~693 MB) and provides a solid basis for streaming speech translation searches.
What is this dataset for?
- Train models capable of translating speech live (streaming speech translation)
- Testing the effectiveness of multilingual transcription/translation systems in real time
- Serve as a base for voice assistants or automatic subtitling applications
Can it be enriched or improved?
Yes. It is possible to increase this dataset with additional languages, audio metadata (duration, accent, speaker) or even annotations on fluidity/translatability. Finer segmentation could also improve performance in streaming mode.
🔎 In summary
🧠 Recommended for
- NLP speech engineers
- Voice translation projects
- Automatic subtitling prototyping
🔧 Compatible tools
- Hugging Face Datasets
- Whisper
- ESPnet
- Fairseq
💡 Tip
For smoother translation, cut long files into shorter speech units.
Frequently Asked Questions
Is this dataset suitable for offline translation?
It was designed for streaming, but can also be used in offline audio translation scenarios.
Can it be easily integrated into a Hugging Face pipeline?
Yes, it is available in Parquet format and compatible with the Hugging Face `datasets` library.
Is it multilingual?
It includes multilingual data, but the set of languages covered depends on the model used for the initial generation.




