---
title: CPT-TrainingData-DPOPairs
canonical_url: "https://www.modelscope.cn/datasets/Tsinghuadhy/CPT-TrainingData-DPOPairs"
md_url: "https://www.modelscope.cn/datasets/Tsinghuadhy/CPT-TrainingData-DPOPairs.md"
repository: Tsinghuadhy/CPT-TrainingData-DPOPairs
chinese_name: "CPT 训练数据 DPO 偏好对 baseline"
last_updated: 2026-05-31
license: MIT
storage_size: "748 MB"
downloads: 4
stars: 0
---

# CPT-TrainingData-DPOPairs

> CPT-TrainingData-DPOPairs - Tsinghuadhy 在 ModelScope 开源的数据集。cpt-trainingdata-dpo-pairs

Tsinghuadhy/CPT-TrainingData-DPOPairs 是 ModelScope 魔搭社区上的数据集，存储大小 748 MB，采用 MIT 许可。

- **Repository**: Tsinghuadhy/CPT-TrainingData-DPOPairs
- **License**: MIT
- **Storage size**: 748 MB
- **Downloads**: 4
- **Stars**: 0
- **Last updated**: 2026-05-31

Source: https://www.modelscope.cn/datasets/Tsinghuadhy/CPT-TrainingData-DPOPairs

---

# cpt-trainingdata-dpo-pairs

Preference-pair data used to train the `DPO+RL` baseline in
[Cognitive Pairwise Training (CPT)](https://github.com/Tsinghua-dhy/CPT) (paper §4.3).

70,352 chosen / rejected preference pairs derived from the same CPT-style trace
pool, labelled by Qwen3-235B-A22B-Instruct-2507.

Train script: [`train/dpo/run_dpo_qwen3_8b.sh`](https://github.com/Tsinghua-dhy/CPT).

## Files
- `train.parquet` — ~744 MB
- `test.parquet` — ~5 MB
