Persona-100k
Massive dataset of 100,000 realistic synthetic profiles, structured to train AI models on personalization, equity, or conversational interaction tasks.
Description
Persona-100k is a textual data set containing 100,000 synthetic profiles, each described by over 70 attributes. These personas cover various aspects such as personal data, profession, cultural preferences, lifestyle, or personality traits. The dataset is entirely in English, in the format .jsonl.
What is this dataset for?
- Train custom LLM models, capable of generating responses adapted to a given user profile
- Evaluate the fairness of models based on simulated demographics
- Simulate realistic users in chatbots or conversational agents
Can it be enriched or improved?
Absolutely. It is possible to add annotation layers (emotional labels, intentions, engagement scores), to adapt personas to other languages, or to enrich each profile with simulated texts (reviews, dialogues, biographies). This dataset is an excellent basis for augmented generation or behavioral alignment.
🔎 In summary
🧠 Recommended for
- AI developers
- Algorithmic equity researchers
- Chatbot designers
🔧 Compatible tools
- Transformers (Hugging Face)
- LangChain
- RAG pipelines
- Recommendation frameworks
💡 Tip
Filter personas by a targeted characteristic (e.g. age, location, interest) to generate rich conditional prompts.
Frequently Asked Questions
Do these profiles correspond to real people?
No, the profiles are 100% synthetic, generated randomly but representative, with no link to real individuals.
Can it be used to test recommendation systems?
Yes, it is an excellent test game to simulate various users in suggestion or customization systems.
Can this dataset be used in a chatbot?
Absolutely. It makes it possible to condition conversational agents to react according to realistic profiles, simulating various scenarios.




