StackLite — Stack Overflow Question Dataset (2016)
This StackLite dataset contains metadata for questions asked on Stack Overflow up to October 2016. Data includes ID, dates, scores, tags, and number of responses. It allows you to analyze trends and correlations between tags and temporal behaviors.
Over 12 million CSV/SQLite questions, with associated metadata
Creative Commons Attribution-ShareAlike (CC BY-SA)
Description
The dataset StackLite offers an excerpt of Stack Overflow questions, focusing on metadata: identifiers, dates, scores, tags, number of answers, and user information. This simplified format facilitates statistical and temporal analysis of programming questions.
What is this dataset for?
- Analyze the evolution of questions by tags over time
- Study the correlations between tags and question scores
- Explore publishing trends by days of the week
Can it be enriched or improved?
Yes, it is possible to add text, questions and answers, or additional user data for finer analyses. A qualitative annotation of questions could also increase its value.
🔎 In summary
🧠 Recommended for
- Data scientists
- Tech trend analysts
- Community Studies Researchers
🔧 Compatible tools
- Pandas
- SQL
- Jupyter
- Table
- PowerBI
💡 Tip
Use in conjunction with the full Stack Exchange dumps for in-depth text analysis.
Frequently Asked Questions
Does this dataset contain the full text of questions and answers?
No, question metadata only. The full text can be found in other larger dumps.
Can this dataset be used to analyze trends in technology?
Yes, it is ideal for observing the popularity and evolution of tags related to programming technologies.
Is this dataset suitable for training an NLP model?
Not directly, as it does not contain textual content, but it can be used as a complement for statistical analyses.




