By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
BillSum - Summaries of American laws
Text

BillSum - Summaries of American laws

The BillSum dataset brings together complete texts of American and California bills with their corresponding summaries, ideal for automatic summarization tasks.

Download dataset
Size

Approximately 23,000 documents, texts and summaries in JSON, total size ~340 MB

Licence

CC0 1.0

Description

‍

The dataset BillSum contains legislative texts from the US Congress and the State of California with summary summaries. Each entry includes the full text of the bill, a summary, and sometimes a title. This corpus is structured to allow the formation of models capable of automatically generating accurate summaries of large legal documents.

‍

‍

What is this dataset for?

‍

  • Train summarization models for legislative texts
  • Facilitate the search and understanding of complex legal documents
  • Develop automated legal assistants for summarizing information

‍

‍

Can it be enriched or improved?

‍

Yes, BillSum can be enriched by adding annotations on the types of laws, temporal metadata, or links to associated court decisions. Multilingual or simplified summaries could also improve its accessibility.

‍

‍

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐⭐ (Ready-to-use data, clear JSON)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low – well-formatted data)
🏷️ Annotation richness⭐⭐⭐✩✩ (Summaries provided, but limited annotations)
📜 Commercial license✅ Yes (CC0 1.0)
👨‍💻 Beginner friendly🌟 Suitable for starting in legal NLP
🔁 Fine-tuning ready🎯 Perfect for fine-tuning on summarization
🌍 Cultural diversity⚠️ Focused on US/California legislation

‍

‍

🧠 Recommended for

  • Legal NLP researchers
  • Automatic summary developers
  • Legal analysts

‍

‍

🔧 Compatible tools

  • Hugging Face Transformers
  • BART
  • Pegasus
  • SpacY

‍

‍

💡 Tip

For best results, clean up long, repetitive sections and test the generation on thematic subsets.

Frequently Asked Questions

Does this dataset only contain American texts?

Yes, it includes bills from the US Congress and the State of California only.

Are summaries generated automatically or manually?

Summaries were created manually, ensuring high quality training.

Can this dataset be used to train multilingual models?

No, the corpus is only in English. However, it is possible to combine it with other datasets for multilingual purposes.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.