BillSum - Summaries of American laws
The BillSum dataset brings together complete texts of American and California bills with their corresponding summaries, ideal for automatic summarization tasks.
Approximately 23,000 documents, texts and summaries in JSON, total size ~340 MB
CC0 1.0
Description
The dataset BillSum contains legislative texts from the US Congress and the State of California with summary summaries. Each entry includes the full text of the bill, a summary, and sometimes a title. This corpus is structured to allow the formation of models capable of automatically generating accurate summaries of large legal documents.
What is this dataset for?
- Train summarization models for legislative texts
- Facilitate the search and understanding of complex legal documents
- Develop automated legal assistants for summarizing information
Can it be enriched or improved?
Yes, BillSum can be enriched by adding annotations on the types of laws, temporal metadata, or links to associated court decisions. Multilingual or simplified summaries could also improve its accessibility.
🔎 In summary
🧠 Recommended for
- Legal NLP researchers
- Automatic summary developers
- Legal analysts
🔧 Compatible tools
- Hugging Face Transformers
- BART
- Pegasus
- SpacY
💡 Tip
For best results, clean up long, repetitive sections and test the generation on thematic subsets.
Frequently Asked Questions
Does this dataset only contain American texts?
Yes, it includes bills from the US Congress and the State of California only.
Are summaries generated automatically or manually?
Summaries were created manually, ensuring high quality training.
Can this dataset be used to train multilingual models?
No, the corpus is only in English. However, it is possible to combine it with other datasets for multilingual purposes.




