---
title: ToxicCommons
canonical_url: "https://www.modelscope.cn/datasets/PleIAs/ToxicCommons"
md_url: "https://www.modelscope.cn/datasets/PleIAs/ToxicCommons.md"
repository: PleIAs/ToxicCommons
last_updated: 2025-06-19
license: "Apache License 2.0"
storage_size: "7.1 GB"
downloads: 590
stars: 0
---

# ToxicCommons

> ToxicCommons - PleIAs 在 ModelScope 开源的数据集。Toxic Commons is a release of 2 million samples of annotated, public domain, multilingual text that was used to train Celadon. It is being released alongside Celadon, in order to better understand multilingual and…

PleIAs/ToxicCommons 是 ModelScope 魔搭社区上的数据集，存储大小 7.1 GB，采用 Apache License 2.0 许可。

- **Repository**: PleIAs/ToxicCommons
- **License**: Apache License 2.0
- **Storage size**: 7.1 GB
- **Downloads**: 590
- **Stars**: 0
- **Last updated**: 2025-06-19

Source: https://www.modelscope.cn/datasets/PleIAs/ToxicCommons

---

# Toxic Commons

Toxic Commons is a release of 2 million samples of annotated, public domain, multilingual text that was used to train [Celadon](https://huggingface.co/PleIAs/celadon). 
It is being released alongside Celadon, in order to better understand multilingual and multicultural toxicity. 

Each sample was classified across 5 axes of toxicity:

*  **Race and origin-based bias**: includes racism as well as bias against someone’s country or region of origin or immigration status, especially immigrant or refugee status. 
*  **Gender and sexuality-based bias**: includes sexism and misogyny, homophobia, transphobia, and sexual harassment. 
*  **Religious bias**: any bias or stereotype based on someone’s religion. 
*  **Ability bias**: bias according to someone’s physical, mental, or intellectual ability or disability. 
*  **Violence and abuse**: overly graphic descriptions of violence, threats of violence, or calls or incitement of violence.


All 2 million samples were classified by a version of Llama 3.1 8B Instruct, with a [custom system prompt](https://github.com/eliotjones1/celadon/blob/main/prompts/annotate.txt).
To replicate the annotation process on your own dataset, feel free to refer to our script [here](https://github.com/eliotjones1/celadon/blob/main/src/2.1_create_annotations.py), and re-create the prompt for your use case. 


Read more about the training details in the paper, [Toxicity of the Commons: Curating Open-Source Pre-Training Data](https://arxiv.org/pdf/2410.22587) by [Catherine Arnett](https://huggingface.co/catherinearnett), [Eliot Jones](https://huggingface.co/eliotj), [Ivan P. Yamshchikov](https://huggingface.co/ivan-the-bearable), [Pierre-Carl Langlais](https://huggingface.co/Pclanglais). 
For more detailed code regarding generating the annotations, please refer to the official [GitHub](https://github.com/Pleias/toxic-commons) repository. 


# How to Cite

```
@article{arnett2024toxicity,
  title={{Toxicity of the Commons: Curating Open-Source Pre-Training Data}},
  author={Arnett, Catherine and Jones, Eliot and Yamshchikov, Ivan P. and Langlais, Pierre-Carl},
  journal={arXiv preprint arXiv:2410.22587},
  url={https://arxiv.org/pdf/2410.22587},
  year={2024}
}
```

# About

Annotations were generated by [Eliot Jones](https://huggingface.co/eliotj) while working at [Pleias](https://huggingface.co/PleIAs). This project was made possible by Jean Zay compute grant #GC011015451.
