UI-Vision GUI Benchmark Dataset
Data set designed for the visual analysis of desktop interfaces, with precise annotations on graphic elements, their function, and their position.
Images + JSON annotations in several subsets (functional, spatial, layout)
MIT
Description
UI-Vision is a visual benchmark for the analysis of desktop graphical interfaces. It contains screenshots annotated in three dimensions: spatial, functional, and structural. Annotations are provided in JSON format, and images are organized by task in dedicated subfolders. This corpus makes it possible to study the identification of UI elements and automated graphical interaction.
What is this dataset for?
- Train models to detect and locate user interface elements
- Evaluate the functional understanding and spatial arrangement of visual components
- Develop AI assistants that can navigate or interact with desktop interfaces
Can it be enriched or improved?
Yes, you can enrich this dataset by adding other types of interfaces (mobile, web), annotation translations, or by simulating real user scenarios. It is also possible to cross-reference this corpus with interaction logs for more advanced behavioral analyses.
🔎 In summary
🧠 Recommended for
- RPA projects
- Smart UI agents
- HMI research
🔧 Compatible tools
- OpenCV
- Detectron2
- Label Studio
- PyTorch
💡 Tip
Use functional and spatial dimensions in synergy to train robust multitasking models.
Frequently Asked Questions
Can this dataset be used to train UI navigation agents?
Yes, it's specifically designed for visual analysis and understanding desktop user interfaces.
What tasks can be evaluated with this corpus?
It makes it possible to assess element detection, functional understanding, and layout analysis in graphical interfaces.
Do the interfaces come from real environments?
Yes, screenshots represent realistic desktop interfaces used in various contexts.




