UW Madison Courses
Text dataset containing grade reports, course lists and sections since 2006, from PDF files extracted by an open source tool.
Over 9,000 courses, 200,000 sections, 3 million notes, SQL files, and PDF snippets
CC0: Public Domain
Description
The dataset UW Madison Courses brings together detailed data on University of Wisconsin-Madison courses, sections, instructors, and grades, covering each semester since 2006. This information comes from official PDF reports that are automatically extracted via an open source tool. This massive corpus makes it possible to analyze academic success and teaching trends.
What is this dataset for?
- Analyzing trends in grades and performances in higher education
- Develop NLP tools to extract and structure educational information
- Build applications to help with course planning or academic prediction
Can it be enriched or improved?
It is possible to add metadata on teachers, courses, or to integrate more recent data. Annotating feelings or levels of difficulty on the courses could enrich the dataset for further analyses.
🔎 In summary
🧠 Recommended for
- Educational data scientists
- NLP researchers
- Academic application developers
🔧 Compatible tools
- SQL
- Pandas
- SpacY
- NLTK
- Jupyter Notebook
💡 Tip
Use this dataset to create predictive models of student success based on courses taken.
Frequently Asked Questions
Does this dataset contain personal information about students?
No, only aggregate grades by section and course are present, no individual personal data.
Can this dataset be used to train a course recommendation model?
Yes, by combining course data, notes and teachers, you can develop personalized recommendation systems.
Is the dataset updated regularly?
The data covers the semesters up to the date of publication, but you should check the source platform for updates.




