By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
SOAR ARC Train 5M — Dataset of 5 million solutions for synthesizing Python programs
Text

SOAR ARC Train 5M — Dataset of 5 million solutions for synthesizing Python programs

SOAR ARC Train 5M is a vast dataset of Python programming solutions linked to the ARC benchmark. It makes it possible to train models capable of self-improvement through iterative learning on successful and failed programs.

Download dataset
Size

Around 5 million examples, solutions in Python code, JSON format or similar

Licence

MIT

Description

SOAR ARC Train 5M is a massive corpus of Python solutions synthesized for the ARC (Abstraction and Reasoning Corpus) benchmark. It includes around 5 million examples, integrating a self-improvement process where failed solutions are reused as valid data through relabelling.

What is this dataset for?

  • Train synthesis models of self-improving programs
  • Test evolutionary learning approaches on complex reasoning tasks
  • Providing a massive corpus for fine-tuning on solving algorithmic problems in Python

Can it be enriched or improved?

This dataset can be extended by adding new programs, relabelling additional synthetic tasks, or integrating annotations on the quality of solutions.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐⭐ (Well-structured data, format suitable for ML workflows)
🧼 Need for cleaning⭐⭐⭐⭐✩ (Low, corpus cleaned by deduplication)
🏷️ Annotation richness⭐⭐⭐⭐✩ (Innovative relabeling for self-learning)
📜 Commercial license✅ Yes (MIT)
👨‍💻 Beginner friendly⚠️ Intermediate to advanced, requires ML and programming skills
🔁 Fine-tuning ready🎯 Perfect for fine-tuning and incremental learning
🌍 Cultural diversity⚠️ Technical specialized dataset, diversity limited to the domain

🧠 Recommended for

  • Researchers in automatic programming
  • ML developers
  • Scalable AI R&D teams

🔧 Compatible tools

  • PyTorch
  • TensorFlow
  • Hugging Face Datasets
  • Python environments

💡 Tip

Use the self-improvement loop to continuously improve your models.

Frequently Asked Questions

What is the main objective of this dataset?

Allow the training of program synthesis models capable of learning from their mistakes and successes.

What is the size and format of the data?

Approximately 5 million Python code examples, often in JSON format or similar.

Is this dataset suitable for beginners?

Rather intended for advanced users, with skills in ML and Python programming.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.