---
title: R1-Distill-Math-Test
canonical_url: "https://www.modelscope.cn/datasets/modelscope/R1-Distill-Math-Test"
md_url: "https://www.modelscope.cn/datasets/modelscope/R1-Distill-Math-Test.md"
repository: modelscope/R1-Distill-Math-Test
chinese_name: "R1蒸馏模型数学推理能力测试集"
last_updated: 2025-08-28
license: "Apache License 2.0"
storage_size: "2.8 MB"
downloads: 7295
stars: 35
---

# R1-Distill-Math-Test

> R1-Distill-Math-Test - modelscope 在 ModelScope 开源的数据集。DeepSeek-R1蒸馏模型数学推理能力测试集

modelscope/R1-Distill-Math-Test 是 ModelScope 魔搭社区上的数据集，存储大小 2.8 MB，采用 Apache License 2.0 许可。

- **Repository**: modelscope/R1-Distill-Math-Test
- **License**: Apache License 2.0
- **Storage size**: 2.8 MB
- **Downloads**: 7295
- **Stars**: 35
- **Last updated**: 2025-08-28

Source: https://www.modelscope.cn/datasets/modelscope/R1-Distill-Math-Test

---

> [!NOTE]
> 该数据集是旧版EvalScope v0.xx使用的数据集，对于 v1.0.0 版本的用户，请使用新版[数据集](https://modelscope.cn/datasets/evalscope/R1-Distill-Math-Test-v2)


# R1蒸馏模型数学推理能力测试集

共728道数学推理题目，包括：
- [MATH-500](https://www.modelscope.cn/datasets/HuggingFaceH4/aime_2024)：一组具有挑战性的高中数学竞赛问题数据集，涵盖七个科目（如初等代数、代数、数论）共500道题。
- [GPQA-Diamond](https://modelscope.cn/datasets/AI-ModelScope/gpqa_diamond/summary)：该数据集包含物理、化学和生物学子领域的硕士水平多项选择题，共198道题。
- [AIME-2024](https://modelscope.cn/datasets/AI-ModelScope/AIME_2024)：美国邀请数学竞赛的数据集，包含30道数学题。

更详细的使用方法详见：[DeepSeek-R1类模型数学能力测试最佳实践](https://evalscope.readthedocs.io/zh-cn/latest/best_practice/deepseek_r1_distill.html)

## 使用方法

### 安装依赖

安装[EvalScope](https://github.com/modelscope/evalscope)模型评估框架：

由于框架在快速迭代中，接口可能不稳定，建议通过源码安装：
```bash
git clone https://github.com/modelscope/evalscope.git
cd evalscope/
pip install -e '.[app]'
```

### 部署模型

使用推理框架部署模型可以加速评测，下面是部署DeepSeek-R1-Distill-Qwen-1.5B模型的示例代码：

**使用vLLM**:

```bash
VLLM_USE_MODELSCOPE=True CUDA_VISIBLE_DEVICES=0 python -m vllm.entrypoints.openai.api_server --model deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B  --served-model-name DeepSeek-R1-Distill-Qwen-1.5B --trust_remote_code --port 8801
```

**使用lmdeploy**：

```bash
LMDEPLOY_USE_MODELSCOPE=True CUDA_VISIBLE_DEVICES=0 lmdeploy serve api_server deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B --model-name DeepSeek-R1-Distill-Qwen-1.5B --server-port 8801
```


### 评测模型

运行下面的python代码评测DeepSeek-R1-Distill-Qwen-1.5B模型在数学推理数据集上的表现：

```python
from evalscope import TaskConfig, run_task
from evalscope.constants import EvalType

task_cfg = TaskConfig(
    model='DeepSeek-R1-Distill-Qwen-1.5B',
    api_url='http://127.0.0.1:8801/v1/chat/completions',
    api_key='EMPTY',
    eval_type=EvalType.SERVICE,
    datasets=[
        'data_collection',
    ],
    dataset_args={
        'data_collection': {
            'dataset_id': 'modelscope/R1-Distill-Math-Test'
        }
    },
    eval_batch_size=64,  # num of workers to seed requests
    generation_config={
        'max_tokens': 20000,  # avoid exceed max length
        'temperature': 0.6,
        'top_p': 0.95,
        'n': 5 # num of repeat for each prompt (note lmdeploy only support n=1)
    },
)

run_task(task_cfg=task_cfg)
```

输出结果：

**这里的计算指标是Pass@1，每个样本重复生成了5次，最终的评测结果是5次的平均值。**

```text
2025-02-10 20:42:56,050 - evalscope - INFO - dataset_level Report:
+-----------+--------------+---------------+-------+
| task_type | dataset_name | average_score | count |
+-----------+--------------+---------------+-------+
|   math    |   math_500   |    0.7832     |  500  |
|   math    |     gpqa     |    0.3434     |  198  |
|   math    |    aime24    |      0.2      |   30  |
+-----------+--------------+---------------+-------+
```

### 结果可视化

EvalScope支持可视化结果，可以查看模型具体的输出。运行以下命令，可以启动可视化界面：
```bash
evalscope app
```

将输出如下链接内容：
```text
* Running on local URL:  http://0.0.0.0:7860
```

点击链接即可看到如下可视化界面，我们需要先选择评测报告然后点击加载：

<p align="center">
  <img src="https://sail-moe.oss-cn-hangzhou.aliyuncs.com/yunlin/images/distill/score.png" alt="alt text" width="100%">
</p>


此外，选择对应的子数据集，我们也可以查看模型的输出内容，观察模型输出是否正确（或者是答案匹配是否存在问题）：

<p align="center">
  <img src="https://sail-moe.oss-cn-hangzhou.aliyuncs.com/yunlin/images/distill/detail.png" alt="alt text" width="100%">
</p>

## 数据集构建方式

使用[EvalScope](https://github.com/modelscope/evalscope)工具构建了该集合，参考[使用教程](https://evalscope.readthedocs.io/zh-cn/latest/advanced_guides/collection/index.html)：

```python
from evalscope.collections import WeightedSampler, CollectionSchema, DatasetInfo
from evalscope.utils.io_utils import dump_jsonl_data

schema = CollectionSchema(name='DeepSeekDistill', datasets=[
            CollectionSchema(name='Math', datasets=[
                    DatasetInfo(name='math_500', weight=1, task_type='math', tags=['en'], args={'few_shot_num': 0}),
                    DatasetInfo(name='gpqa', weight=1, task_type='math', tags=['en'],  args={'subset_list': ['gpqa_diamond'], 'few_shot_num': 0}),
                    DatasetInfo(name='aime24', weight=1, task_type='math', tags=['en'], args={'few_shot_num': 0}),
            ])
        ])

#  get the mixed data
mixed_data = WeightedSampler(schema).sample(100000)  # set a large number to ensure all datasets are sampled
dump_jsonl_data(mixed_data, 'test.jsonl')
```


## 下载方法 
:modelscope-code[]{type="sdk"}
:modelscope-code[]{type="git"}
