PhilPapers Papers Summarized Labeled
Text dataset containing 3,776 philosophical papers with summaries generated by LLM and multi-label classification in 17 philosophical schools.
3,776 documents in JSON format, including titles, summaries, multi-label labels
MIT
Description
PhilPapers Papers Summarized Labeled is a dataset composed of 3,776 philosophical documents from PhilPapers. Each document contains the title, a summary summary of 2-3 sentences generated by an LLM, and a multi-label classification into 17 different philosophical schools, such as Existentialism, Stoicism, or Utilitarianism.
What is this dataset for?
- Train multi-label classification models on academic texts
- Develop automatic summary tools in the philosophical field
- Analyze and explore trends in schools of thought through documents
Can it be enriched or improved?
Yes, it is possible to add more languages, improve summaries with newer models, or complete annotations with finer subcategories.
🔎 In summary
🧠 Recommended for
- NLP researchers
- Philosophy students
- Academic tool developers
🔧 Compatible tools
- Hugging Face Datasets
- Scikit-learn
- PyTorch
- Transformers
💡 Tip
Use multi-label classification to train models that can handle complex and multiple themes in a single document.
Frequently Asked Questions
What is the format of the summaries provided in this dataset?
Summaries are short sentences (2-3 sentences) generated automatically by an LLM.
Can this dataset be used for multi-label classification?
Yes, it contains multi-label labels for 17 philosophical schools.
Can other philosophical categories be added to the dataset?
Yes, the dataset can be enriched with additional annotations or more accurate subcategories.




