StreamVLN Trajectory Data
Dataset containing visual observations (RGB images) and action annotations collected in a simulator for vision-language navigation tasks in Matterport3D.
Several thousand RGB images and JSON trajectory annotations, spread over several sub-datasets
CC BY-SA 4.0
Description
StreamVLN Trajectory Data combines RGB images and detailed annotations of actions in simulated Matterport3D environments. It combines several open-source Vision-and-Language navigation datasets, making it easy to learn trajectories guided by textual instructions.
What is this dataset for?
- Training autonomous navigation models in vision-language
- Evaluate agents who can follow instructions in simulated environments
- Develop multimodal systems combining vision and language for robotics
Can it be enriched or improved?
This dataset can be enriched by adding videos, trajectories, finer annotations on the environment, or additional sensorimotor data from the simulator.
🔎 In summary
🧠 Recommended for
- Robotics researchers
- Multimodal AI developers
- Simulation engineers
🔧 Compatible tools
- Habitat simulator
- PyTorch
- TensorFlow
- VLN frameworks
💡 Tip
Use annotations to train agents to precisely follow instructions in simulated 3D environments.
Frequently Asked Questions
What data is included in this dataset?
Annotated RGB images and action sequences in Matterport3D simulated environments for navigation.
Can this dataset be used for autonomous navigation in the real world?
Indirectly, it is mainly used to train and evaluate models in simulation, but can be used as a basis for transfer to reality.
How are annotations structured?
The annotations are in JSON and describe the navigation instructions and the sequence of discrete actions corresponding to the images.




