By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
IRC Disentangle
Text

IRC Disentangle

This dataset provides manually annotated IRC logs with response structures, allowing you to reconstruct tangled conversations. A valuable resource for training or evaluating models for understanding dialogue.

Download dataset
Size

77,563 JSON-annotated IRC messages with response graphs

Licence

CC-BY 4.0

Description

‍

The dataset IRC Disentangle offers over 77,000 messages from IRC channels, each annotated manually to indicate the response links between messages. It thus makes it possible to reconstruct mixed conversations typical of online chats. Each message comes with a raw version, an ASCII version, a tokenization, and response graphs showing the associated messages. The game is divided into channels and dates, with particular attention paid to the quality of the annotations.

‍

‍

What is this dataset for?

‍

  • Train conversational deinterlacing models (disentanglement)
  • Test a chatbot's ability to identify threads in a disorganized flow
  • Improve tools for moderating and analysing online dialogues

‍

‍

Can it be enriched or improved?

‍

Yes, the dataset can be enriched by adding other IRC channels, translating to other languages, or adding additional annotations (emotions, intentions, etc.). Graphs can also be used to generate realistic synthetic conversations, useful for fine-tuning dialogue models.

‍

‍

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐✩✩ (Clear structure but requires familiarity with graphs)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low – data already normalized)
🏷️ Annotation richness⭐⭐⭐⭐✩ (Very good – explicit connections between messages)
📜 Commercial license✅ Yes (CC-BY 4.0)
👨‍💻 Beginner friendly⚠️ Accessible but requires understanding of graphs
🔁 Fine-tuning ready🎯 Yes, for dialogue disentanglement or multi-turn chatbots
🌍 Cultural diversity⚠️ Medium – English only (Ubuntu logs)

‍

‍

🧠 Recommended for

  • Researchers in NLP dialogue
  • AI moderators
  • Conversation reenactment projects

‍

‍

🔧 Compatible tools

  • PyTorch
  • Hugging Face Datasets
  • NetworkX
  • Transformers

‍

‍

💡 Tip

Use response graphs to automatically segment multi-threaded dialogues before fine-tuning a conversational model.

Frequently Asked Questions

Does this dataset contain full dialogs or only isolated messages?

It contains messages linked together via annotation graphs, allowing complete conversations to be reconstructed.

Can it be used to train a multi-user chatbot?

Yes, it's ideal for training a model that can handle multiple conversations embedded in the same flow.

Are there tools provided to process this dataset?

Yes, scripts are available in the associated GitHub repository for data conversion, evaluation, and preprocessing.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.