MMRL
Data set designed to train multimodal models in reinforcement learning. Combines text, images, and CoT thought chains.
Multimodal data (text + image), optimized for RL training, structured format for fine-tuning
Apache 2.0
Description
MMRL is a data set designed to train multimodal reasoning models through various stages of reinforcement learning (RL). It aligns logical reasoning and visual comprehension using specific rewards related to the conciseness, structure, and quality of the thought chains generated. This dataset is used in training the Revisual-R1 (7B) model.
What is this dataset for?
- Train multimodal models to generate reasoning from images and text
- Optimizing RL agents for complex visual-textual tasks
- Test advanced strategies like PAD and reward by effective length
Can it be enriched or improved?
Yes. The dataset can be enriched by adding new image-text pairs or by annotating finer arguments. It is also possible to adjust reward signals to test other learning criteria.
🔎 In summary
🧠 Recommended for
- Researchers in multimodal RL
- Visual-textual LLM developers
- Advanced AI laboratories
🔧 Compatible tools
- Transformers
- TRL
- DeepSpeed
- PyTorch
- Hugging Face vision tools
💡 Tip
To maximize the effectiveness of rewards, test different weightings between conciseness, logic and vision during training.
Frequently Asked Questions
What does MMRL offer compared to a classic RL dataset?
It integrates a multimodal dimension combining images and text, with rewards adapted to the reasoning and the CoT structure.
Is the dataset suitable for training small models?
It is optimized for medium to large models, but can be sampled for lighter tasks.
Does MMRL include vision-specific annotations?
Yes, visual data is aligned with textual reasoning, allowing deep training on the image-logical link.



