HaGrid — Hand Gesture Recognition Image Dataset
Massive dataset of hand gesture images (552,992 samples), divided into 18 classes, with accurate annotations for detection and tracking.
552,992 Full HD images annotated in COCO + 21 landmark points, 18 classes
CC BY-SA 4.0
Description
HaGrid is a vast hand-gesture recognition dataset, composed of more than 550,000 Full HD images captured under various conditions (natural light, artificial light, backlight, etc.). It covers 18 gesture classes as well as a “no_gesture” class for noise. The images are accompanied by COCO annotations (bounding boxes), 21 landmarks, and additional information on the main hand.
What is this dataset for?
- Train gesture recognition models for video conferencing (Zoom, Skype...)
- Develop gestural interfaces in the automotive sector or home automation
- Test algorithms for detecting the hand and tracking precise movements
Can it be enriched or improved?
Yes, the dataset can be adapted to specific use cases by filtering certain classes or by supplementing with videos. It is also possible to improve diversity by adding other populations or cultural contexts. Annotations can also be refined for multi-hand segmentation or classification tasks.
🔎 In summary
🧠 Recommended for
- Computer vision researchers
- HCI developers
- Gestural interaction projects
🔧 Compatible tools
- Detectron2
- YoloV5
- MediaPipe
- Tensorflow Object Detection API
💡 Tip
Use user identifiers to avoid data leakage between train and test during cross-validation.
Frequently Asked Questions
Does this dataset contain videos or only images?
It only contains Full HD images, but with a wide range of poses and gestures that simulate a sequence.
Are the gestures annotated accurately?
Yes, each hand is annotated with a bounding box, 21 landmarks, and metadata about the main hand and trust.
Is it suitable for training models in real time?
Yes, by extracting subsets or resizing images, it can be used for embedded or real-time applications.




