By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
7K Books with Metadata
Text

7K Books with Metadata

This dataset includes more than 6,500 books with their metadata (title, author, description, ISBN, etc.), obtained via the Google Books API from Goodreads data. It is ideal for training a recommendation engine, clustering or applying NLP techniques (literary text analysis, genre classification, etc.).

Download dataset
Size

6542 lines in CSV format, containing titles, authors, authors, descriptions, ISBNs, language, categories

Licence

CC0: Public Domain

Description

The dataset 7K Books with Metadata was built using the Google Books API and Goodreads. It contains over 6,500 books, each with a complete sheet: title, author (s), description, ISBN, language, categories, and other useful attributes. It is particularly suitable for work in automatic language processing (NLP) and in data science applied to literature or publishing.

What is this dataset for?

  • Create a book recommendation engine (content-based filtering)
  • Analyze literary trends by categories or descriptions
  • Experimenting with NLP: automatic summary, gender classification, semantic clustering

Can it be enriched or improved?

Yes. It is possible to complete the metadata with other sources (Amazon, OpenLibrary), to extract keywords from the summaries, or to filter by language to refine the NLP models. The author himself plans to expand the dataset with future larger API requests.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐⭐ (Ready-to-use CSV format)
🧼 Need for cleaning⭐⭐⭐⭐☆ (Low: may contain a few null values or duplicates)
🏷️ Annotation richness⭐⭐⭐⭐⭐ (High: multiple structured and textual fields — descriptions, categories, etc.)
📜 Commercial license✅ Yes (CC0)
👨‍💻 Beginner friendly📚 Great dataset to start with recommendation or NLP tasks
🔁 Fine-tuning ready🔥 Good support for fine-tuning (useful for recommendation or NLP tasks thanks to textual fields)
🌍 Cultural diversity🌍 Moderate: mainly in English but some descriptions specify a language

🧠 Recommended for

  • Data science students
  • Literature lovers
  • Literary AI researchers

🔧 Compatible tools

  • Pandas
  • Scikit-learn
  • SpacY
  • HuggingFace Transformers

💡 Tip

Use TF-IDF analysis of abstracts to generate thematic clusters or launch an internal book search engine.

Frequently Asked Questions

Does this dataset contain the full texts of the books?

No, only the metadata is present (title, author, summary, categories...), but not the full content of the works.

Can we filter by language or by category?

Yes, some columns allow you to know the language and genres associated with each book.

Can a recommendation engine be trained with this dataset?

Yes, it is even one of its main uses. The “description” field allows filtering based on content or similarity.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.