---
title: Llama3-Chinese-dataset
canonical_url: "https://www.modelscope.cn/datasets/baicai003/Llama3-Chinese-dataset"
md_url: "https://www.modelscope.cn/datasets/baicai003/Llama3-Chinese-dataset.md"
repository: baicai003/Llama3-Chinese-dataset
chinese_name: "Llama3 中文化数据集"
last_updated: 2024-05-11
license: "Apache License 2.0"
storage_size: "4.0 GB"
domain:
  - configs
tasks:
  - "{config_name=sft_zh_with_all, data_files=[{split=train, path=sft_zh_with_all.jsonl}]}"
downloads: 11774
stars: 102
---

# Llama3-Chinese-dataset

> Llama3-Chinese-dataset - baicai003 在 ModelScope 开源的数据集。Llama3 中文化数据集

baicai003/Llama3-Chinese-dataset 是 ModelScope 魔搭社区上的{config_name=sft_zh_with_all, data_files=[{split=train, path=sft_zh_with_all.jsonl}]}数据集，涉及 configs 领域，存储大小 4.0 GB，采用 Apache License 2.0 许可。

- **Repository**: baicai003/Llama3-Chinese-dataset
- **License**: Apache License 2.0
- **Tasks**: {config_name=sft_zh_with_all, data_files=[{split=train, path=sft_zh_with_all.jsonl}]}
- **Domain**: configs
- **Storage size**: 4.0 GB
- **Downloads**: 11774
- **Stars**: 102
- **Last updated**: 2024-05-11

Source: https://www.modelscope.cn/datasets/baicai003/Llama3-Chinese-dataset

---

### 介绍
Github: https://github.com/CrazyBoyM/llama3-Chinese-chat  
该数据集已统一处理为firefly格式，可以配合firefly工具直接训练llama3中文模型。  
特别推荐：新增DPO偏好对齐中文强化学习数据：https://modelscope.cn/datasets/shareAI/shareAI-Llama3-DPO-zh-en-emoji/summary

### 使用方法
本数据集可以在firefly框架上直接加载使用训练
#### 直接使用
sft_zh_with_all.jsonl文件是包含所有清洗处理后数据集的合并文件，可以直接使用FireFly训练你的中文模型。（过滤后问答数据量约169万条）
#### 按需使用
下面演示如何合并以及处理数据，你可以按照自己的需求对单点数据进行加强扩充。  
首先，将所有数据集合并为一个jsonl文件：
```shell
cat *.jsonl > merged.jsonl
```
然后，使用check_jsonl.py对数据集进行校验, 去除格式错误的个别样本：
```python
python3 check_jsonl.py
```
最后，使用change_info.py修改模型的身份认知信息：
```python
python3 change_info.py
```
小建议：
- 如果觉得个别数据集质量不高，可以自行删除。
- 如果觉得个别数据集样本不够，可以自行添加。
- 如果觉得个别数据集样本太多，可以自行采样。
附如何随机采样一部分：
```python
import json
import random

def sample_jsonl(input_file, output_file):
    # 读取输入文件
    with open(input_file, 'r', encoding='utf-8') as f:
        lines = f.readlines()

    # 随机采样1/100的数据
    sample_size = len(lines) // 100
    sampled_lines = random.sample(lines, sample_size)

    # 写入输出文件
    with open(output_file, 'w', encoding='utf-8') as f:
        for line in sampled_lines:
            f.write(line)

if __name__ == "__main__":
    input_file = "./1_0_firefly_chinese_common_task_1649k.jsonl"  # 输入文件路径
    output_file = "sampled_chinese_common_task_16k.jsonl"  # 输出文件路径
    sample_jsonl(input_file, output_file)
```

### 下载方法 
:modelscope-code[]{type="sdk"}
:modelscope-code[]{type="git"}
