LEMONADE — Egocentric Video Language Assessment
LEMONADE is a benchmark composed of closed-ended questions associated with egocentric video clips recorded in the kitchen, making it possible to assess behavioral understanding, temporal reasoning and biomechanical analysis.
Description
The dataset LEMONADE includes 36,521 closed-ended questions linked to self-centered videos from the EPFL-Smart-Kitchen-30 benchmark. The questions cover three main categories: behavioral understanding, long-term understanding, and biomechanical analysis via 3D pose estimation.
What is this dataset for?
- Evaluate video comprehension of multimodal language models
- Test temporal reasoning and inference on long video sequences
- Analyzing biomechanical data from video (postures, movements)
Can it be enriched or improved?
Yes. It is possible to consider adding other modalities such as sensor data, fine annotations of behavior, or rich text transcripts. Integration with other video benchmarks can also enrich the ecosystem.
🔎 In summary
🧠 Recommended for
- Vision-language researchers
- Multimodal Video Model Developers
- Biomechanics R&D teams
🔧 Compatible tools
- PyTorch
- Hugging Face Transformers
- MMAction2
- OpenPose
- Multimodal notebooks
💡 Tip
Combine this dataset with additional behavioral annotations to enrich contextual understanding.
Frequently Asked Questions
What are the main question categories in LEMONADE?
The questions are classified into three broad categories: behavioral understanding, long-term understanding, and biomechanical analysis.
Can this dataset be used to train a multimodal model?
Yes, it is suitable for training and evaluating models combining video and language.
What data modalities are present in this dataset?
The dataset includes egocentric videos in MP4 format and associated questions and answers as a CSV file.




