---
title: ChatParts_Dataset
canonical_url: "https://www.modelscope.cn/datasets/shellwork/ChatParts_Dataset"
md_url: "https://www.modelscope.cn/datasets/shellwork/ChatParts_Dataset.md"
repository: shellwork/ChatParts_Dataset
last_updated: 2024-09-26
license: "Apache License 2.0"
storage_size: "554 MB"
domain:
  - text
tasks:
  - iGEM
downloads: 25
stars: 0
---

# ChatParts_Dataset

> ChatParts_Dataset - shellwork 在 ModelScope 开源的数据集。📚 Dataset Information

shellwork/ChatParts_Dataset 是 ModelScope 魔搭社区上的iGEM数据集，涉及 text 领域，存储大小 554 MB，采用 Apache License 2.0 许可。

- **Repository**: shellwork/ChatParts_Dataset
- **License**: Apache License 2.0
- **Tasks**: iGEM
- **Domain**: text
- **Storage size**: 554 MB
- **Downloads**: 25
- **Stars**: 0
- **Last updated**: 2024-09-26

Source: https://www.modelscope.cn/datasets/shellwork/ChatParts_Dataset

---

## 📚 Dataset Information

This dataset is utilized for fine-tuning the following models:

- [shellwork/ChatParts-llama3.1-8b](https://www.modelscope.cn/datasets/shellwork/ChatParts-llama3.1-8b)
- [shellwork/ChatParts-qwen2.5-14b](https://www.modelscope.cn/datasets/shellwork/ChatParts-qwen2.5-14b)

### 📁 File Structure

The dataset is organized as follows:

```plaintext
D:\ChatParts_Dataset
│
├── README.md
├── Original_data
│   ├── iGEM_competition_web.rar
│   ├── paper_txt_processed.rar
│   └── wiki_data.rar
└── Training_dataset
    ├── pt_txt.json
    ├── sft_eval.json
    └── sft_train.json
```
- **Original_data:**
  - `iGEM_competition_web.rar`: Contains raw text documents scraped from iGEM competition websites.
  - `paper_txt_processed.rar`: Contains processed text from over 1,000 synthetic biology review papers.
  - `wiki_data.rar`: Contains raw Wikipedia data related to synthetic biology.

  The original data was collected using web crawlers and subsequently filtered and manually curated to ensure quality. These raw `.txt` documents serve as the foundational learning passages for the model's pre-training phase. The consolidated and processed text can be found in the `pt_txt.json` file within the `Training_dataset` directory.

- **Training_dataset:**
  - `pt_txt.json`: Consolidated and preprocessed text passages used for the model's pre-training step.
  - `sft_train.json`: Contains over 180,000 question-answer pairs derived from the original documents, used for supervised fine-tuning (SFT) training.
  - `sft_eval.json`: Contains over 20,000 question-answer pairs reserved for evaluating the model post-training, maintaining a 9:1 data ratio compared to the training set.

  The `sft_train.json` and `sft_eval.json` files consist of meticulously organized question-answer pairs extracted from all available information in the original documents. These datasets facilitate the model's supervised instruction learning process, enabling it to generate accurate and contextually relevant responses.

### 📄 License

This dataset is released under the **Apache License 2.0**. For more details, please refer to the [license information](https://github.com/shellwork/XJTLU-Software-RAG/tree/main) in the repository.

## 🔗 Additional Resources

- **RAG Software:** Explore the full capabilities of our Retrieval-Augmented Generation software [here](https://github.com/shellwork/XJTLU-Software-RAG/tree/main).
- **Training Data:** Access and review the extensive training dataset [here](https://www.modelscope.cn/datasets/shellwork/ChatParts_Dataset).


---

Feel free to reach out through our GitHub repository for any questions, issues, or contributions related to this dataset.
