LLM Prompt Recovery Data Gemini and Gemma
This dataset contains pairs of original and reworded prompts, as well as texts generated by two systems: Gemini and Gemma. It is intended to improve and assess the ability of LLM models to understand and reformulate prompts.
Several hundreds of text examples in JSON/CSV format (not specified)
Apache 2.0
Description
LLM Prompt Recovery Data Gemini and Gemma gathers data on the original test prompts, the prompts provided to Gemini and Gemma, and the rewritten texts generated by these two models. It is a useful corpus for the analysis and improvement of the capacities for reformulation and generation of prompts.
What is this dataset for?
- Evaluate the quality of reformulation of prompts using LLM models
- Train models to understand and rewrite text instructions
- Improve the automatic generation of prompts for specific tasks
Can it be enriched or improved?
This dataset can be supplemented with qualitative annotations on the consistency or fidelity of the reformulations, as well as with additional data from other models to increase diversity.
🔎 In summary
🧠 Recommended for
- NLP researchers
- LLM developers
- Data scientists
🔧 Compatible tools
- Hugging Face Transformers
- TensorFlow
- PyTorch
- Jupyter Notebooks
💡 Tip
Check the consistency of the reformulations before training to guarantee the quality of the fine-tuning.
Frequently Asked Questions
What type of data does this dataset contain?
Original prompts, reformulated prompts and texts generated by two models named Gemini and Gemma.
Is this dataset suitable for training reformulation models?
Yes, it contains pairs that are useful for training and evaluating the reformulation of prompts.
Does the license allow commercial use?
Yes, the Apache 2.0 license allows commercial and modified use.




