SVLA SO100 Sorting Dataset
Multimodal dataset combining video sequences and detailed robotic data from an SO100 robot. It includes 30fps HD videos as well as robot action and status data, making it possible to study and model the manipulation of objects by vision and commands.
18 videos (480x640, 30fps, AV1 codec), 6633 frames, action and observation data in parquet format
Apache 2.0
Description
The dataset SVLA SO100 Sorting Dataset contains 18 HD videos (480x640, 30 fps) recorded with an SO100 robot, accompanied by action (six degrees of freedom) and observation data. The data is organized into episodes and chunks with detailed metadata.
What is this dataset for?
- Develop and train vision-language-action models in robotics.
- Analyze robotic movements and interactions for sorting and handling tasks.
- Test multimodal learning algorithms integrating video and sensor data.
Can it be enriched or improved?
This dataset can be enriched by adding new video sequences, additional annotations on actions, or extensions to other types of robots or tasks. The accuracy of the status data can also be increased by fine annotation.
🔎 In summary
🧠 Recommended for
- Robotics researchers
- Multimodal AI developers
- Vision-action R&D teams
🔧 Compatible tools
- Video frameworks (OpenCV, FFmpeg)
- Parquet tools
- Robotic libraries (ROS)
💡 Tip
Synchronize video and sensor data well for effective multimodal training.
Frequently Asked Questions
What is the resolution and format of the videos in the dataset?
The videos are 480x640 pixels, 30 frames per second, encoded in AV1 codec without audio.
What action data is provided with the videos?
Six degrees of freedom are captured: shoulder, elbow, wrist movements, and grip opening.
Is this dataset suitable for supervised learning or reinforcement?
Primarily for multimodal supervised learning, but can be adapted for reinforcement with additional data.




