By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
40k Songs with Audio Features and Lyrics
Multimodal

40k Songs with Audio Features and Lyrics

Multimodal dataset combining song lyrics and extracted audio characteristics, covering 79 musical genres and approximately 43,000 titles, ideal for music analysis projects and multimodal learning.

Download dataset
Size

Around 43,000 songs in English with lyrics files (text) and audio features (CSV/JSON)

Licence

Apache 2.0

Description

The dataset 40k Songs with Audio Features and Lyrics brings together approximately 43,000 songs in English, including lyrics and explicit audio characteristics from three sources combined. It covers 79 varied musical genres, allowing a rich exploration of multimodal musical data.

What is this dataset for?

  • Musical analysis and classification based on lyrics and audio characteristics
  • Training multimodal models combining audio and text
  • Search by automatically generating music or lyrics

Can it be enriched or improved?

The dataset can be supplemented with additional annotations such as mood, tempo, or more accurate metadata. The integration of complete raw audio data would also broaden its use.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐✩✩ (Tabular data easy to handle, requires text and audio processing)
🧼 Need for cleaning⭐⭐⭐⭐✩ (Moderate – standardization of lyrics and audio features needed)
🏷️ Annotation richness⭐⭐⭐✩✩ (Good balance between text and audio, annotations limited to extracted features)
📜 Commercial license✅ Yes (Apache 2.0)
👨‍💻 Beginner friendly⚠️ Moderate, requires audio and NLP knowledge
🔁 Fine-tuning ready🎵 Very suitable for multimodal fine-tuning
🌍 Cultural diversity🎭 Good diversity thanks to 79 music genres

🧠 Recommended for

  • AI music researchers
  • Multimodal projects
  • NLP + audio developers

🔧 Compatible tools

  • Librosa
  • Hugging Face Transformers
  • PyTorch
  • TensorFlow
  • SpacY

💡 Tip

Carefully pretreating speech (cleaning and tokenization) optimizes the performance of multimodal models.

Frequently Asked Questions

What languages are present in this dataset?

Mostly English, with a wide range of musical genres.

Can we use this dataset to generate music?

Yes, especially for training multimodal models combining audio and text.

Does the Apache 2.0 license allow commercial use?

Yes, this license is permissive and allows free commercial use.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.