By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
LLM Prompt Recovery Data Gemini and Gemma
Text

LLM Prompt Recovery Data Gemini and Gemma

This dataset contains pairs of original and reworded prompts, as well as texts generated by two systems: Gemini and Gemma. It is intended to improve and assess the ability of LLM models to understand and reformulate prompts.

Download dataset
Size

Several hundreds of text examples in JSON/CSV format (not specified)

Licence

Apache 2.0

Description

LLM Prompt Recovery Data Gemini and Gemma gathers data on the original test prompts, the prompts provided to Gemini and Gemma, and the rewritten texts generated by these two models. It is a useful corpus for the analysis and improvement of the capacities for reformulation and generation of prompts.

What is this dataset for?

  • Evaluate the quality of reformulation of prompts using LLM models
  • Train models to understand and rewrite text instructions
  • Improve the automatic generation of prompts for specific tasks

Can it be enriched or improved?

This dataset can be supplemented with qualitative annotations on the consistency or fidelity of the reformulations, as well as with additional data from other models to increase diversity.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐✩ (Clear structure, textual data easy to handle)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low, well-formatted data)
🏷️ Annotation richness⭐⭐⭐✩✩ (Medium, prompts and outputs but no human evaluations)
📜 Commercial license✅ Yes (Apache 2.0)
👨‍💻 Beginner friendly⚠️ Suitable for beginner NLP researchers
🔁 Fine-tuning ready⚠️ Useful for fine-tuning on paraphrasing and generation
🌍 Cultural diversity🇬🇧 Neutral, textual data in English

🧠 Recommended for

  • NLP researchers
  • LLM developers
  • Data scientists

🔧 Compatible tools

  • Hugging Face Transformers
  • TensorFlow
  • PyTorch
  • Jupyter Notebooks

💡 Tip

Check the consistency of the reformulations before training to guarantee the quality of the fine-tuning.

Frequently Asked Questions

What type of data does this dataset contain?

Original prompts, reformulated prompts and texts generated by two models named Gemini and Gemma.

Is this dataset suitable for training reformulation models?

Yes, it contains pairs that are useful for training and evaluating the reformulation of prompts.

Does the license allow commercial use?

Yes, the Apache 2.0 license allows commercial and modified use.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.