---
title: Fineweb-Edu-Chinese-V3
canonical_url: "https://www.modelscope.cn/datasets/opencsg/Fineweb-Edu-Chinese-V3"
md_url: "https://www.modelscope.cn/datasets/opencsg/Fineweb-Edu-Chinese-V3.md"
repository: opencsg/Fineweb-Edu-Chinese-V3
last_updated: 2026-09-08
license: other
storage_size: "21 MB"
downloads: 198
stars: 0
---

# Fineweb-Edu-Chinese-V3

> Fineweb-Edu-Chinese-V3 - opencsg 在 ModelScope 开源的数据集。Fineweb-Edu-Chinese-V3（抽样版）

opencsg/Fineweb-Edu-Chinese-V3 是 ModelScope 魔搭社区上的数据集，存储大小 21 MB，采用 other 许可。

- **Repository**: opencsg/Fineweb-Edu-Chinese-V3
- **License**: other
- **Storage size**: 21 MB
- **Downloads**: 198
- **Stars**: 0
- **Last updated**: 2026-09-08

Source: https://www.modelscope.cn/datasets/opencsg/Fineweb-Edu-Chinese-V3

---

# Fineweb-Edu-Chinese-V3（抽样版）

<div align="center">
  <a href="#chinese">中文</a> | <a href="#english">English</a>
</div>

<div align="center">

[OpenCSG 社区](https://opencsg.com/datasets) | [数据集许可协议](./OpenCSG数据集许可协议.md)

</div>

> ### ⚠️ 本仓库为抽样版本
>
> 本仓库收录 Fineweb-Edu-Chinese-V3 的**抽样子集**，用于快速预览与评估：**7,729 篇文档 / 12,677 条样本 / 21.1 MB**，约为完整版的 **10%（按文档）。
>
> **完整版请见 Hugging Face：**
> **https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V3**
>
> 抽样按文档为单位随机选取（固定随机种子，可复现），同一文档的样本不会被拆散。除规模外，抽样版与完整版在构建管线、数据格式、字段定义和许可条款上完全一致。

<a id="chinese"></a>

## 数据集简介

**Fineweb-Edu-Chinese-V3** 是 OpenCSG 面向学科知识问答、教材理解和推理型指令微调场景构建的高质量中英双语教育 SFT 数据集，也是 Fineweb-Edu-Chinese 系列的最新版本。

完整版包含 **18.81 万条 SFT 样本**，来自 **100,442 篇**高质量图书、教材、学科文献与技术长文，覆盖计算机、自然科学、社科人文、法学、经济五大学科方向，并同步提供 **Messages**、**Messages-no-system**、**Alpaca** 三种训练格式。三种格式是同一批问答对的不同导出视图，训练时应按模型模板选择其中一种，而不是简单相加作为独立数据规模。

V3 是 Fineweb-Edu-Chinese 系列的一次**数据源与构造范式的整体切换**。V1.0 至 V2.3 均以大规模中文网页语料为基础：通过打分器筛选出具备教育属性的网页文本，再由大模型生成问答。V3 不再从网页出发，而是直接以**正式出版的图书、教材与学科文献**为源，每一条样本都由一条 **8 阶段文档理解管线**逐本构建：先做版面解析与章节重建，再抽取"知识点"（Knowledge Point）作为最小可测试单元，然后经过问题草拟、原文证据检索、难度精炼、跨章节合并、题型转换，最后导出为训练格式。

这条链路的目标不是产出更多样本，而是让每条问答都能追溯到源文档中的具体章节与实体，并覆盖完整的推导过程而非孤立的概念复述。

---

## 核心价值

面向学科知识的中文 SFT 数据长期存在三个实际问题：高质量教材与专著难以转化为可训练格式；网页合成问答缺乏可追溯的证据支撑；生成结果偏向浅层概念复述，缺少推导过程与结构化表达。

本数据集重点提升三类能力：

- **教材级知识密度**：数据源为正式出版的图书、教材与学科文献，而非网页抓取内容，知识点具备体系性与准确性基础。
- **有据可依的问答构造**：问题与答案由源文档中显式抽取的实体（定理、关键公式、命题、表格）驱动，构建阶段引入原文证据检索环节，降低自由发挥式回答的比例。
- **推理型长答案**：答案普遍包含分步推导、符号与单位定义、假设与适用范围说明，适合训练需要展示完整解题过程的模型。

---

## 版本演进

| 版本号 | 核心定位 | 数据规模 | 关键特性与改进 | 当前状态 |
| --- | --- | --- | --- | --- |
| **V1.0** | 概念验证 | 约 9000 万条，约 300GB | 初代 Chinese Fineweb Edu 语料；BERT 打分模型；MinHash 去重；数据源包括 CCI2、SkyPile、Tele-AI | 已弃用 |
| **V2.0** | 规模化扩展 | 约 1.88 亿条，约 420B tokens | 升级至 OpenCSG csg-wukong-enterprise V2 打分器；扩展 Industry2、wanjuan1.0、wudao 等数据源 | 已弃用 |
| **V2.1** | 预训练精选 | 总计约 1.5T tokens | 按分数分层组织；新增 map-cc、opencsg-cc；支持灵活预训练和课程学习 | 推荐用于预训练 |
| **V2.2** | SFT 与对齐 | 约 143.7 万条高质量问答 | 将高质量教育语料转化为 SFT 问答数据；提供纯 QA 与上下文版本 | 历史 SFT 版本 |
| **V2.3** | 更高纯度的 SFT 数据 | 23.04 万条 QA pairs | 升级 V2.2 的源文本选择和生成逻辑；强化证据对齐、质量过滤和多格式导出 | 历史 SFT 版本 |
| **V3** | **文档级学科推理数据** | **18.81 万条样本 / 10.04 万篇源文档** | **数据源由网页语料切换为正式出版图书与教材；8 阶段文档理解管线；知识点驱动构造；新增选择题与表格题型；单样本文本量提升至 V2.3 的 2.6 倍** | **推荐用于 SFT** |

> **关于跨版本规模的可比性**：V1.0 至 V2.1 是**预训练语料**，规模以条数、GB 或 tokens 计量；V2.2 起系列转向 **SFT 问答数据**，规模以 QA pairs 计量。因此表中"约 9000 万条"与"18.81 万条"并非同一量纲，跨阶段直接比较条数没有意义。**仅 V2.2、V2.3、V3 之间的样本数具有可比性。**

系列的演进可以分为两个阶段：

- **预训练语料阶段（V1.0 → V2.1）**：目标是从海量中文 Web 内容中筛出更具教育价值的文本。改进集中在打分器（BERT → csg-wukong-enterprise V2）、数据源扩展与分数分层组织上。
- **SFT 数据阶段（V2.2 → V3）**：目标转为构造可直接训练的问答数据。V2.2 完成了从语料到问答的形态转换；V2.3 收紧了源文本筛选门槛；**V3 则更换了数据源本身**——从中文网页语料切换为正式出版的图书与教材。

换句话说，V1.0 至 V2.3 是同一条技术路线的逐步收紧，数据源始终是网页语料，改进集中在"如何筛得更准"；V3 改变的是这条路线的起点，重心从"筛选"转向"文档理解"。

### 系列定位与生态

Fineweb-Edu-Chinese 系列是全球下载量排名前三的中文数据集之一，累计下载超百万次，已在学术与产业两端形成规模化使用：

- **学术**：被斯坦福大学、清华大学、中国人民大学高瓴人工智能学院、上海人工智能实验室、北京智源研究院等 20 余家机构的论文引用，累计被 100 余篇学术论文引用，出现在 NeurIPS、ACL、EMNLP、ICLR 等国际会议及 Nature 子刊、JMLR 等期刊中；合作机构还包括鹏城实验室、西南电子技术研究所、西班牙国家级超算中心（Barcelona Supercomputing Center）与 Mozilla Data Collective。
- **产业**：支撑 Llama3-Chinese、DeepSeek 等模型训练，并被中国移动、中国联通、英伟达（NVIDIA）、苹果公司（Apple Inc.）、OPPO、美团、阿里巴巴、蚂蚁集团、面壁智能（ModelBest）、Krafton 等企业采用。
- **生态**：系列累计数据体量达 2.42TB、覆盖 9.57 亿条高质量文本，已孵化出 10 余个垂直领域微调模型。

> 以上为 Fineweb-Edu-Chinese **系列**截至 V2.3 的累计生态数据，用于说明本系列的定位与沿革，并非 V3 单个版本的使用统计。

系列的一贯理念是让中文大模型不只是"读到更多中文"，而是能够"学到更好的中文"。V3 在此基础上再进一步：不只是"学到更好的中文"，而是**学会完整的推导与解题过程**。

---

## V3 相比 V2.3 的变化

### 数据来源与构造范式

V2.3 的核心挑战是"如何从海量网页中筛出适合生成 SFT 样本的文本"——它训练了一个中文源文本分类打分器，从约 2.3T 语料中排序选择。这条路线的上限受制于网页本身：教育类网页的知识往往是碎片化的、缺少推导过程的，且与文章排版、导航栏、广告混杂。

V3 直接绕开了这个问题：源文档本身就是成体系的图书与教材，知识密度和准确性由出版流程保证。管线的重点也随之从"筛选"转向"理解"——如何把一本 PDF 教材正确地解析为章节结构、识别其中的定理与公式、并围绕它们构造出能覆盖完整推导链的问题。

### 关键指标对比

| 维度 | V2.3 | V3（完整版） |
| --- | --- | --- |
| 数据来源 | 中文网页语料（约 2.3T 候选） | 正式出版图书、教材、学科文献、技术长文 |
| 源单位 | 网页片段 | **100,442 篇完整文档** |
| 筛选/构造核心 | Yuan-embedding 打分器筛选 + GPT-4.1 mini 生成 | **8 阶段文档理解管线 + 知识点驱动构造** |
| 样本数 | 230,400 | 188,148 |
| 单样本平均体量 | 1,691 字节 | **4,384 字节（2.6×）** |
| 文本总量（Alpaca 格式） | 约 390 MB | **约 825 MB（2.1×）** |
| 题型 | General QA | **General QA / Table QA / Single Choice / Multiple Choice** |
| 答案风格 | 段落式解释 | **分步推导，含符号定义、单位、假设与适用范围** |
| 训练格式 | Messages / Messages-no-sys / Alpaca | 同左 |
| 语言 | 中文为主 | 中英双语 |

### 如何理解规模变化

样本条数在本系列中已连续两个版本下降：V2.2 的 143.7 万 → V2.3 的 23.04 万 → V3 的 18.81 万。这是系列的一贯取向——V2.3 发布时即说明，其规模小于 V2.2「并不是因为数据能力下降，而是因为筛选标准更严格」。V3 延续了同一逻辑，但下降幅度小得多（-18%），且伴随单样本体量的大幅上升。

V3 的单条样本平均为 4,384 字节，是 V2.3（1,691 字节）的 **2.6 倍**，因此整体文本量反而是 V2.3 的 **2.1 倍**。换言之，V3 的条数下降并不意味着数据量减少，而是同样的文本预算被分配到了更少、更长、信息更完整的样本上。

差异来自样本形态：V2.3 的问答多为概念解释型的段落式回答；V3 的问题通常包含背景设定、符号与单位表格、多个分小问，答案则是分步推导过程。实测 V3 样本的问题平均 2,320 字符、答案平均 1,722 字符，P90 分别达到 4,321 和 3,343 字符。

因此，**如果你的目标是训练模型输出简洁的知识性回答，V2.3 仍然是合适的选择；如果目标是让模型展示完整的解题与推导过程，V3 更契合**。两者数据源不重叠（网页语料 vs. 出版文献），语言侧重也不同（V2.3 以中文为主，V3 为中英双语），因此完全可以混合使用以兼顾两类能力。

### 新增能力

- **客观题型**：V3 新增 Single Choice 与 Multiple Choice，合计占完整版的 28.4%，可直接用于训练或评测选择题作答能力，这是 V2.3 不具备的。
- **表格理解**：Table QA 来自源文档中的真实表格，而非合成数据。
- **文档级溯源**：每条样本携带 `source_paper` 与所属章节信息，可回溯到具体出版物，便于做数据审计与领域切分。
- **学科可控性**：16 个构建域按学科独立配置抽取范式与人设，支持按学科方向做采样或分层训练。

---

## 数据构建流程

原始文档经过 8 个阶段的管线处理，逐本构建为 SFT 样本。各阶段职责如下：

| 阶段 | 名称 | 职责 |
| --- | --- | --- |
| **S1** | Preprocess | 版面解析与清洗，将 PDF/文档转为结构化 Markdown，分离正文、公式、表格与图片 |
| **S1.5** | ChunkConnect | 按章节标题重建文档层级，将正文切分为带 `heading` 的语义块 |
| **S2** | Editor | 抽取知识点（KP）：识别定理、关键公式、命题、表格等实体，产出带 `section_heading`、`original_entity`、`knowledge_point_summary`、`question_types` 的原子/宏观知识点 |
| **S3** | Query Generator | 基于知识点与源文档生成问题草稿 `query_draft` |
| **S4** | Retriever | 回到原文检索支撑证据，为每个问题补齐 `core_answer` 与 `background_evidence` |
| **S5** | Refiner | 提升问题难度与完整性，产出结构化的 `question` / `final_answer` |
| **S6** | Merger | 跨章节合并知识点，筛选并生成综合性"超级问题"，附带质量评分 |
| **S7** | Converter | 转换为最终题型：General QA / Table QA / Single Choice / Multiple Choice |
| **S8** | SFT Formatter | 按文档切分 train/val，导出 Messages、Messages-no-system、Alpaca 三种格式 |

其中 **S2 的知识点抽取**是整条链路的核心设计。它遵循"整体完整性原则"（Holistic Integrity Principle）：当一个命题连同它的前置假设与关键公式共同构成一个自洽的推导闭环时，会被合并为单个**宏观知识点**而不被拆散；不构成这种闭环的实体则作为**原子知识点**处理。这一设计使得下游问题能够覆盖完整的推导链，而不是退化为孤立的公式辨认题。

管线为每个学科方向配置了独立的抽取范式、few-shot 示例与领域人设 system prompt，共覆盖 16 个构建域。

---

## 数据规格与仓库组织

本仓库以 **16 个 `.tar.gz` 抽样包**组织，每个包对应一个数据来源。解压后顶层为**文档目录**，每个文档目录内包含该文档的全部训练格式与统计文件。

| 文件 | 说明 |
| --- | --- |
| `train_messages.jsonl` / `val_messages.jsonl` | Messages 格式，含领域人设 system prompt |
| `train_messages_no_sys.jsonl` / `val_messages_no_sys.jsonl` | Messages 格式，不含 system |
| `train_alpaca.jsonl` / `val_alpaca.jsonl` | Alpaca 格式 |
| `stats.json` | 该文档的样本统计 |



---

## 数据统计（抽样版）




### 题型分布（抽样版）

| 题型 | 样本数 | 占比 | 完整版占比 |
| --- | --- | --- | --- |
| General QA | 9,763 | 77.0% | 69.5% |
| Single Choice | 1,470 | 11.6% | 15.6% |
| Multiple Choice | 1,157 | 9.1% | 12.8% |
| Table QA | 287 | 2.3% | 2.1% |



---

## Schema

### Messages

```json
{
  "messages": [
    {
      "role": "system",
      "content": "You are an expert engineering educator with deep knowledge of mechanics of materials, structural mechanics, electrical circuit analysis, sensors and instrumentation, mechanical design, and applied mathematics for engineering. Provide rigorous, step-by-step answers grounded in physical principles and engineering design standards."
    },
    { "role": "user", "content": "用户问题" },
    { "role": "assistant", "content": "助手答案" }
  ],
  "metadata": {
    "domain": "cbs",
    "question_type": "General QA",
    "source_paper": "29838-设计构成",
    "quality_score": 0.0,
    "images": []
  }
}
```

### Alpaca

```json
{
  "instruction": "Answer the following cbs question with detailed derivations and explanations.",
  "input": "用户问题",
  "output": "助手答案",
  "metadata": { "...": "同上" }
}
```

### metadata 字段

| 字段 | 含义 |
| --- | --- |
| `domain` | 构建域标识，共 16 个（如 `disciplinary_cs`、`highqua1_na`、`cbs`、`csdn_article_b1`） |
| `question_type` | 题型：`General QA` / `Table QA` / `Single Choice` / `Multiple Choice` |
| `source_paper` | 源文档标识 |
| `quality_score` | 质量分字段。当前公开快照中该字段普遍为 `0.0`，质量控制主要发生在 S6 筛选阶段而非记录级打分，**不建议将其作为过滤依据** |
| `images` | 源文档中与该样本所属章节关联的图片路径。**图片文件本身不包含在本仓库中**，该字段仅作溯源标记，本数据集为纯文本 SFT 数据 |

---

## 真实样例

**Question**

> **Background**
>
> Engineering problem. The relationship between contrast and harmony is fundamental in color design, influencing both emotional expression and visual stability.
>
> **Variables and parameters.**
>
> | Symbol | Definition | Unit |
> |---|---|---|
> | C | Contrast level | dimensionless |
> | H | Harmony level | dimensionless |
> | E | Emotional expression | dimensionless |

**Answer**

> 1\) (a) Contrast and harmony are interdependent aspects in color design. Contrasting colors enhance emotional expression, often evoking excitement and vibrancy in a visual composition.
>
> 2\) (b) Harmonious arrangements provide stability and unity in visual communication by reducing the intensity of contrasts, creating a cohesive aesthetic.
>
> 3\) (c) To achieve a balance between contrast and harmony, a designer could use complementary colors for focal points while employing analogous colors for background elements.

该样例体现了典型结构：问题包含背景设定、符号与单位表格、分小问设问；答案按 `1) (a)`、`2) (b)` 的编号逐问作答。

---

## 快速开始

### 下载与解压

```bash
# 通过 ModelScope CLI 下载抽样版
pip install -U modelscope
modelscope download --dataset opencsg/Fineweb-Edu-Chinese-V3 --local_dir ./Fineweb-Edu-Chinese-V3

# 或使用 git（需要 git-lfs）
git lfs install
git clone https://www.modelscope.cn/datasets/opencsg/Fineweb-Edu-Chinese-V3.git

# 解压单个包
tar -xzf 北京大学出版社电子教材_sft_sample10pct.tar.gz -C ./data/
```

Python 方式：

```python
from modelscope.hub.snapshot_download import snapshot_download

snapshot_download(
    "opencsg/Fineweb-Edu-Chinese-V3",
    repo_type="dataset",
    local_dir="./Fineweb-Edu-Chinese-V3",
)
```

### 加载数据

```python
import glob
from datasets import load_dataset

# 递归 glob 兼容单层与三层两种目录布局
train_files = glob.glob("./data/**/train_messages.jsonl", recursive=True)
val_files   = glob.glob("./data/**/val_messages.jsonl",   recursive=True)

ds = load_dataset("json", data_files={"train": train_files, "validation": val_files})
print(ds["train"][0]["messages"])
```

### 重新切分（推荐）

由于原始切分按文档进行、train/val 比例偏离常规，建议合并后自行切分：

```python
files = glob.glob("./data/**/*_messages.jsonl", recursive=True)
full = load_dataset("json", data_files=files, split="train")
split = full.train_test_split(test_size=0.02, seed=42)
```

### 获取完整版

抽样版适合快速验证数据格式与质量。正式训练请使用完整版：

```bash
pip install -U huggingface_hub
hf download opencsg/Fineweb-Edu-Chinese-V3 --repo-type dataset --local-dir ./Fineweb-Edu-Chinese-V3-full
```

### 三种格式的选择

- **Messages（含 system）**：带学科人设 system prompt，推荐用于需要稳定领域角色的 chat-style SFT。
- **Messages-no-system**：不带 system 的纯对话，适合已有自定义 system prompt 的训练流程。
- **Alpaca**：`instruction` / `input` / `output` 结构，兼容 LLaMA-Factory 等 Alpaca 系训练框架。

---

## 适用场景

- 学科知识问答与教育助手模型的监督微调
- 需要展示完整推导过程的推理型模型训练
- 中英双语学科问答能力构建
- 教材理解、专业培训与科研辅助类模型
- 文档理解管线、知识点抽取与合成数据质量研究

> 本抽样版更适合**数据格式验证、质量抽查与流程调试**。正式训练建议使用完整版。

---

## 限制与风险边界

本数据集为基于真实出版文档、通过大模型管线构造的**合成问答数据**。尽管构建过程中引入了原文证据检索与质量筛选环节，样本仍可能包含事实错误、推导缺陷、遗漏或表达偏差，**不应被视为事实权威来源或专业意见**。

使用前请特别注意以下几点：

- **本仓库仅为抽样子集**：覆盖完整版约 6.7% 的样本，不适合直接用于正式训练。
- **抽样分布与完整版存在偏差**：`csdn_article` 抽样率偏低（4.8%），导致抽样版中选择题占比低于完整版。
- **切分比例非标准**：抽样版全局 train/val 约为 39 : 61，直接使用会导致训练集偏小，建议重新切分（见上文）。
- **`quality_score` 不可用于过滤**：该字段在当前快照中普遍为 `0.0`。
- **图片不随仓库分发**：`metadata.images` 中的路径指向构建时的中间目录，仅作溯源用途，无法直接访问。
- **语言分布不均**：system prompt 与相当比例的问答内容为英文，中文内容主要集中在源自中文教材与技术长文的包中。

将本数据集用于医疗、法律、金融、教育评价等高风险场景前，需进行额外的领域专家评审、模型评测、安全评测与合规审查。

---

## 许可说明

使用本数据集需要遵循 [OpenCSG 数据集许可协议](./OpenCSG数据集许可协议.md)。仓库 metadata 中的 `license: other` 表示本数据集采用平台预设列表之外的许可协议，实际许可条款以该协议为准。

本数据集可按 OpenCSG 数据集许可协议申请商业用途。若计划将本数据集，或基于本数据集训练、增强的模型、系统、Agent、API 服务和商业产品用于商业场景，请发送邮件至 lorraineg@opencsg.com 获取许可。

---

## Citation

```bibtex
@dataset{opencsg_fineweb_edu_chinese_v3_2026,
  title        = {Fineweb-Edu-Chinese-V3: A Document-Grounded Bilingual Educational Instruction Dataset},
  author       = {OpenCSG},
  year         = {2026},
  url          = {https://www.modelscope.cn/datasets/opencsg/Fineweb-Edu-Chinese-V3},
  note         = {OpenCSG dataset repository}
}
```

---
---

<a id="english"></a>

# Fineweb-Edu-Chinese-V3 (Sample)

> ### ⚠️ This repository is a sampled subset
>
> It contains a **sampled subset** of Fineweb-Edu-Chinese-V3 for quick preview and evaluation: **7,729 documents / 12,677 samples / 21.1 MB**, roughly **10%** of the full release.
>
> **For the full dataset, see Hugging Face:**
> **https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V3**
>
> Sampling is done at the document level with a fixed seed (reproducible); samples from one document are never split apart. Apart from scale, the sampled version is identical to the full release in construction pipeline, data format, field definitions, and license terms.

## Dataset Overview

**Fineweb-Edu-Chinese-V3** is a high-quality bilingual (Chinese/English) educational SFT dataset built by OpenCSG for disciplinary knowledge QA, textbook comprehension, and reasoning-oriented instruction tuning. It is the latest release in the Fineweb-Edu-Chinese series.

The full release contains **188,148 SFT samples** derived from **100,442 documents** — books, textbooks, disciplinary literature, and long-form technical articles — spanning computer science, natural sciences, social sciences and humanities, law, and economics, exported in **Messages**, **Messages-no-system**, and **Alpaca** formats. These are alternative views of one QA set; select the format matching your model template rather than summing them as independent data volume.

V3 marks a **shift in both data source and construction paradigm** for the series. V1.0 through V2.3 were built on large-scale Chinese web corpora: a scorer selected web text with educational properties, and an LLM then generated QA from it. V3 starts from **formally published books, textbooks, and disciplinary literature** instead, constructing every sample document-by-document through an **8-stage document understanding pipeline**.

---

## Core Value

- **Textbook-grade knowledge density**: sources are formally published books, textbooks, and disciplinary literature rather than scraped web pages.
- **Evidence-grounded QA construction**: questions and answers are driven by entities explicitly extracted from source documents (theorems, key equations, propositions, tables), with a source-retrieval stage that reduces free-form hallucinated answers.
- **Reasoning-style long answers**: answers typically include step-by-step derivations, symbol and unit definitions, and statements of assumptions and validity ranges.

---

## Version Evolution

| Version | Positioning | Scale | Key Features and Improvements | Status |
| --- | --- | --- | --- | --- |
| **V1.0** | Proof of concept | ~90M records, ~300GB | First-generation Chinese Fineweb Edu corpus; BERT scorer; MinHash dedup; sources include CCI2, SkyPile, Tele-AI | Deprecated |
| **V2.0** | Scale-up | ~188M records, ~420B tokens | Upgraded to OpenCSG csg-wukong-enterprise V2 scorer; added Industry2, wanjuan1.0, wudao | Deprecated |
| **V2.1** | Pretraining selection | ~1.5T tokens total | Score-stratified organization; added map-cc, opencsg-cc; supports flexible pretraining and curriculum learning | Recommended for pretraining |
| **V2.2** | SFT and alignment | ~1.437M QA pairs | Converted high-quality educational corpus into SFT QA data; pure-QA and with-context variants | Legacy SFT release |
| **V2.3** | Higher-purity SFT data | 230.4K QA pairs | Upgraded V2.2's source selection and generation logic; strengthened evidence alignment, quality filtering, multi-format export | Legacy SFT release |
| **V3** | **Document-level disciplinary reasoning data** | **188.1K samples / 100.4K source documents** | **Source switched from web corpora to published books and textbooks; 8-stage document understanding pipeline; knowledge-point-driven construction; adds multiple-choice and table question types; 2.6× larger per-sample text than V2.3** | **Recommended for SFT** |

> **On cross-version comparability**: V1.0 through V2.1 are **pretraining corpora**, measured in records, GB, or tokens; from V2.2 onward the series shifted to **SFT QA data**, measured in QA pairs. "~90M records" and "188.1K samples" are not the same unit. **Only V2.2, V2.3, and V3 have directly comparable sample counts.**

The series evolved in two stages:

- **Pretraining corpus stage (V1.0 → V2.1)**: selecting educationally valuable text from large-scale Chinese web content.
- **SFT data stage (V2.2 → V3)**: constructing directly trainable QA data. V2.2 completed the corpus-to-QA transformation; V2.3 tightened source-selection thresholds; **V3 replaced the data source itself**.

### Series Positioning and Ecosystem

The Fineweb-Edu-Chinese series is among the top three most-downloaded Chinese datasets worldwide, with over one million cumulative downloads:

- **Academia**: cited in papers from over 20 institutions including Stanford University, Tsinghua University, Renmin University's Gaoling School of Artificial Intelligence, Shanghai AI Laboratory, and BAAI; cited by 100+ academic papers across NeurIPS, ACL, EMNLP, ICLR, Nature-family journals, and JMLR.
- **Industry**: supports the training of models such as Llama3-Chinese and DeepSeek, adopted by China Mobile, China Unicom, NVIDIA, Apple Inc., OPPO, Meituan, Alibaba, Ant Group, ModelBest, and Krafton.
- **Ecosystem**: 2.42TB of cumulative data covering 957 million high-quality texts, with 10+ vertical-domain fine-tuned models incubated on top of it.

> These are cumulative ecosystem statistics for the Fineweb-Edu-Chinese **series** as of V2.3, not usage statistics for V3 alone.

---

## What Changed in V3 vs. V2.3

| Dimension | V2.3 | V3 (full) |
| --- | --- | --- |
| Data source | Chinese web corpora (~2.3T candidates) | Published books, textbooks, disciplinary literature, technical articles |
| Source unit | Web page fragment | **100,442 complete documents** |
| Core of selection/construction | Yuan-embedding scorer + GPT-4.1 mini generation | **8-stage document understanding pipeline + knowledge-point-driven construction** |
| Samples | 230,400 | 188,148 |
| Avg. bytes per sample | 1,691 | **4,384 (2.6×)** |
| Total text (Alpaca format) | ~390 MB | **~825 MB (2.1×)** |
| Question types | General QA | **General QA / Table QA / Single Choice / Multiple Choice** |
| Answer style | Paragraph-form explanation | **Step-by-step derivation with symbol definitions, units, assumptions** |
| Language | Primarily Chinese | Bilingual (Chinese/English) |

### Understanding the Scale Change

Sample counts have declined across two consecutive releases: 1.437M (V2.2) → 230.4K (V2.3) → 188.1K (V3). When V2.3 was released, its smaller size was explained as "not a decline in data capability, but the result of stricter selection criteria." V3 follows the same logic with a far smaller decline (-18%) and a substantial increase in per-sample volume.

V3 averages 4,384 bytes per sample versus V2.3's 1,691 — a **2.6× increase** — so total text volume is **2.1× that of V2.3**. Measured over V3 samples, questions average 2,320 characters and answers 1,722 characters, with P90 values of 4,321 and 3,343.

**If your goal is concise knowledge-style answers, V2.3 remains appropriate; if the goal is complete derivation and problem-solving processes, V3 fits better.** The two do not overlap in source data and can be mixed.

---

## Statistics (Sampled Version)


### Question Type Distribution

| Question Type | Samples | Share | Full-release share |
| --- | --- | --- | --- |
| General QA | 9,763 | 77.0% | 69.5% |
| Single Choice | 1,470 | 11.6% | 15.6% |
| Multiple Choice | 1,157 | 9.1% | 12.8% |
| Table QA | 287 | 2.3% | 2.1% |


---

## Quick Start

```bash
pip install -U modelscope
modelscope download --dataset opencsg/Fineweb-Edu-Chinese-V3 --local_dir ./Fineweb-Edu-Chinese-V3

tar -xzf 北京大学出版社电子教材_sft_sample10pct.tar.gz -C ./data/
```

```python
import glob
from datasets import load_dataset

train_files = glob.glob("./data/**/train_messages.jsonl", recursive=True)
val_files   = glob.glob("./data/**/val_messages.jsonl",   recursive=True)
ds = load_dataset("json", data_files={"train": train_files, "validation": val_files})
```

For the full release:

```bash
pip install -U huggingface_hub
hf download opencsg/Fineweb-Edu-Chinese-V3 --repo-type dataset --local-dir ./Fineweb-Edu-Chinese-V3-full
```

---

## Limitations and Risk Boundaries

This is **synthetic QA data** constructed from real published documents via an LLM pipeline. Samples may contain factual errors, flawed derivations, omissions, or biased phrasing, and **should not be treated as an authoritative source**.

- **This repository is a sampled subset** covering ~6.7% of the full release; it is not suitable for production training on its own.
- **Sampling distribution deviates from the full release**: `csdn_article` was sampled at a lower rate (4.8%), lowering the multiple-choice share.
- **Non-standard split ratio**: the sampled version's global train/val ratio is about 39 : 61; re-splitting is recommended.
- **`quality_score` is not usable as a filter**: generally `0.0` in this snapshot.
- **Images are not distributed**: paths under `metadata.images` are provenance markers only.
- **Uneven language distribution**: system prompts and a substantial share of QA content are in English.

Before use in high-risk scenarios such as healthcare, law, finance, or educational assessment, conduct additional domain-expert review, model evaluation, safety evaluation, and compliance review.

---

## License

Use of this dataset is subject to the [OpenCSG Dataset License Agreement](./OpenCSG数据集许可协议.md). The `license: other` field indicates a license outside the platform's preset list; the agreement itself governs the actual terms.

Commercial use may be requested under the OpenCSG Dataset License Agreement. Please email lorraineg@opencsg.com to obtain a license.
