By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
SVLA SO100 Sorting Dataset
Video

SVLA SO100 Sorting Dataset

Multimodal dataset combining video sequences and detailed robotic data from an SO100 robot. It includes 30fps HD videos as well as robot action and status data, making it possible to study and model the manipulation of objects by vision and commands.

Download dataset
Size

18 videos (480x640, 30fps, AV1 codec), 6633 frames, action and observation data in parquet format

Licence

Apache 2.0

Description

The dataset SVLA SO100 Sorting Dataset contains 18 HD videos (480x640, 30 fps) recorded with an SO100 robot, accompanied by action (six degrees of freedom) and observation data. The data is organized into episodes and chunks with detailed metadata.

What is this dataset for?

  • Develop and train vision-language-action models in robotics.
  • Analyze robotic movements and interactions for sorting and handling tasks.
  • Test multimodal learning algorithms integrating video and sensor data.

Can it be enriched or improved?

This dataset can be enriched by adding new video sequences, additional annotations on actions, or extensions to other types of robots or tasks. The accuracy of the status data can also be increased by fine annotation.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐✩✩ (Requires video processing and Parquet sensor data)
🧼 Need for cleaning⭐⭐⭐⭐✩ (Moderate – structured format but multiformat to manage)
🏷️ Annotation richness⭐⭐⭐⭐✩ (Good – detailed action data and metadata)
📜 Commercial license✅ Yes (Apache 2.0)
👨‍💻 Beginner friendly⚠️ Moderate – video and sensor formats need mastery
🔁 Fine-tuning ready🎯 Suitable for multimodal learning in robotics
🌍 Cultural diversity⚖️ Technical content without notable cultural bias

🧠 Recommended for

  • Robotics researchers
  • Multimodal AI developers
  • Vision-action R&D teams

🔧 Compatible tools

  • Video frameworks (OpenCV, FFmpeg)
  • Parquet tools
  • Robotic libraries (ROS)

💡 Tip

Synchronize video and sensor data well for effective multimodal training.

Frequently Asked Questions

What is the resolution and format of the videos in the dataset?

The videos are 480x640 pixels, 30 frames per second, encoded in AV1 codec without audio.

What action data is provided with the videos?

Six degrees of freedom are captured: shoulder, elbow, wrist movements, and grip opening.

Is this dataset suitable for supervised learning or reinforcement?

Primarily for multimodal supervised learning, but can be adapted for reinforcement with additional data.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.