---
title: Magpie-Pro-DPO-200K
canonical_url: "https://www.modelscope.cn/datasets/Magpie-Align/Magpie-Pro-DPO-200K"
md_url: "https://www.modelscope.cn/datasets/Magpie-Align/Magpie-Pro-DPO-200K.md"
repository: Magpie-Align/Magpie-Pro-DPO-200K
last_updated: 2025-01-16
license: "Apache License 2.0"
storage_size: "695 MB"
downloads: 1367
stars: 1
---

# Magpie-Pro-DPO-200K

> Magpie-Pro-DPO-200K - Magpie-Align 在 ModelScope 开源的数据集。This dataset is still under internal assessment. Please use it with caution!

Magpie-Align/Magpie-Pro-DPO-200K 是 ModelScope 魔搭社区上的数据集，存储大小 695 MB，采用 Apache License 2.0 许可。

- **Repository**: Magpie-Align/Magpie-Pro-DPO-200K
- **License**: Apache License 2.0
- **Storage size**: 695 MB
- **Downloads**: 1367
- **Stars**: 1
- **Last updated**: 2025-01-16

Source: https://www.modelscope.cn/datasets/Magpie-Align/Magpie-Pro-DPO-200K

---

This dataset is still under internal assessment. Please use it with caution!

To create this dataset, we carefully selected a diverse range of high-quality instructions from Magpie datasets, with a particular emphasis on Math and Coding tasks. We then generate responses from the Llama-3 base model using URIAL as rejected. Then, we generate responses from Qwen2-72B-Instruct and Llama-3-8B-Instruct and take the instruction-response pair as chosen.

### Other Magpie DPO Datasets
We observed that the following DPO datasets may have better performance after we burned a lot of GPU hours :)

|Model Name | Dataset | Type | Description |
|-------------|:-------|:-------|:-------|
| [Llama 3 8B Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct) | [Magpie-Air-DPO-100K](https://huggingface.co/datasets/Magpie-Align/Magpie-Air-DPO-100K-v0.1) | DPO | DPO dataset via Best-of-N sampling and rewards.
| [Llama 3 70B Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct) | [Magpie-Pro-DPO-100K](https://huggingface.co/datasets/Magpie-Align/Magpie-Pro-DPO-100K-v0.1) | DPO | DPO dataset via Best-of-N sampling and rewards.
| [Llama 3.1 70B Instruct](https://huggingface.co/meta-llama/Meta-Llama-3.1-70B-Instruct) | [Magpie-Llama-3.1-Pro-DPO-100K](https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-DPO-100K-v0.1) | DPO | DPO dataset via Best-of-N sampling and rewards.
