Letterbox Movie Classification Dataset
A structured data set including 10,002 movies, with complete metadata (titles, genres, directors, duration, language, etc.) and textual descriptions usable in NLP.
10,002 movies, CSV file with 15 columns (texts, categories, numeric)
CC0: Public Domain
Description
The Letterbox Movie Classification Dataset contains the metadata of more than 10,000 movies, including information such as title, genres, genres, language, duration, director, and user engagement metrics (likes, watches, lists). Each movie also comes with a text description, ideal for NLP applications.
What is this dataset for?
- Create movie recommendation systems based on content or popularity
- Analyze genre, studio, or rating trends over time
- Carry out the classification or clustering of movies according to user behaviors or themes
Can it be enriched or improved?
Yes, the dataset can be enriched by links to trailers, images, or even by a more advanced thematic classification. It is also possible to improve text columns via embeddings or semantic analyses.
🔎 In summary
🧠 Recommended for
- Cinema Data Analysts
- Recommendation engine developers
- NLP students
🔧 Compatible tools
- Pandas
- Scikit-learn
- TensorFlow
- Hugging Face Transformers
💡 Tip
Combine text descriptions with genres and engagement metrics to build a more accurate hybrid recommendation engine.
Frequently Asked Questions
Are movie descriptions usable for NLP?
Yes, each movie has a text description, which makes it possible to apply tasks such as feeling analysis or topic modeling.
Can it be used for a personalized recommendation engine?
Absolutely, the numerous columns allow for both collaborative and content-based approaches.
Does the dataset only cover recent movies?
No, it includes films from a variety of periods, languages, and genres, covering a wide range of cinematographic diversity.




