---
title: reasoning-v1-20m
canonical_url: "https://www.modelscope.cn/datasets/AI-ModelScope/reasoning-v1-20m"
md_url: "https://www.modelscope.cn/datasets/AI-ModelScope/reasoning-v1-20m.md"
repository: AI-ModelScope/reasoning-v1-20m
last_updated: 2025-03-20
license: "Apache License 2.0"
storage_size: "81 GB"
downloads: 2748
stars: 1
---

# reasoning-v1-20m

> reasoning-v1-20m - AI-ModelScope 在 ModelScope 开源的数据集。We are excited to release a synthetic reasoning dataset containing 22mil+ general reasoning questions and responses generated using deepseek-ai/DeepSeek-R1-Distill-Llama-70B. While there have been multiple…

AI-ModelScope/reasoning-v1-20m 是 ModelScope 魔搭社区上的数据集，存储大小 81 GB，采用 Apache License 2.0 许可。

- **Repository**: AI-ModelScope/reasoning-v1-20m
- **License**: Apache License 2.0
- **Storage size**: 81 GB
- **Downloads**: 2748
- **Stars**: 1
- **Last updated**: 2025-03-20

Source: https://www.modelscope.cn/datasets/AI-ModelScope/reasoning-v1-20m

---

![image/png](https://cdn-uploads.huggingface.co/production/uploads/637d41b2bb031d2afee723ae/v0tl4UPoPIp-d0mQGLlX6.png)

We are excited to release a synthetic reasoning dataset containing 22mil+ general reasoning questions and responses generated using [deepseek-ai/DeepSeek-R1-Distill-Llama-70B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B). While there have been multiple efforts to build open reasoning datasets for math and code tasks, we noticed a lack of large datasets containing reasoning traces for diverse non code/math topics like social and natural sciences, education, creative writing and general conversations, which is why we decided to release this dataset.<br> 
*Note: Please note that in this instance we have not verified the reasoning traces and answers for accuracy.*


**Dataset details:**<br>
*Total number of rows:* 22.2 million rows<br>
*Total number of tokens:* 35.8 billion tokens

The dataset can be used to fine-tune smaller, more efficient models to mimic the reasoning capabilities of larger models like DeepSeek-R1 using SFT.

**Response format:**
```
<think>
-- reasoning trace --
</think>
-- answer --
```

**Loading the dataset:**
```python
from datasets import load_dataset
ds = load_dataset("glaiveai/reasoning-v1-20m", split="train")
```
