---
title: OpenThoughts-114k-math
canonical_url: "https://www.modelscope.cn/datasets/open-r1/OpenThoughts-114k-math"
md_url: "https://www.modelscope.cn/datasets/open-r1/OpenThoughts-114k-math.md"
repository: open-r1/OpenThoughts-114k-math
last_updated: 2025-02-11
license: "Apache License 2.0"
storage_size: "935 MB"
downloads: 5757
stars: 1
---

# OpenThoughts-114k-math

> OpenThoughts-114k-math - open-r1 在 ModelScope 开源的数据集。This is a filtered and metadata enriched version of open-thoughts/OpenThoughts-114k.

open-r1/OpenThoughts-114k-math 是 ModelScope 魔搭社区上的数据集，存储大小 935 MB，采用 Apache License 2.0 许可。

- **Repository**: open-r1/OpenThoughts-114k-math
- **License**: Apache License 2.0
- **Storage size**: 935 MB
- **Downloads**: 5757
- **Stars**: 1
- **Last updated**: 2025-02-11

Source: https://www.modelscope.cn/datasets/open-r1/OpenThoughts-114k-math

---

This is a filtered and metadata enriched version of [`open-thoughts/OpenThoughts-114k`](https://huggingface.co/datasets/open-thoughts/OpenThoughts-114k).

While the original dataset is a valuable resource containing [DeepSeek-R1](https://huggingface.co/deepseek-ai/DeepSeek-R1) outputs, it has very little metadata (only 2 fields: `system` and `conversations`). It does not contain, for instance, the original solution label, which means that we can not verify the model answers.

## What we did
- filtered the dataset for math content (math questions were prefixed by "Return your final response within \\boxed{}." -- see [here](https://github.com/open-thoughts/open-thoughts/blob/main/open_thoughts/math/reason.py#L16C43-L16C90))
- found the original questions in the [`AI-MO/NuminaMath-CoT`](https://huggingface.co/datasets/AI-MO/NuminaMath-CoT) and mapped them back to each generation
- verified model generations using our [Math-Verify library](https://github.com/huggingface/Math-Verify)
- added a metadata field with the token count of each DeepSeek-R1 completion

## Data structure
- `source`: original `source` from Numina-Math
- `problem`: problem statement, from Numina-Math
- `solution`: original solution/gold label, from Numina-Math
- `messages`: message turns for finetuning on the correct solutions, from Numina-Math
- `system`: system prompt sent to DeepSeek-R1, from OpenThoughts
- `conversations`: message turns from the DeepSeek-R1 generation. The last turn is the model output, from OpenThoughts
- `generated_token_count`: number of tokens (counted using the DeepSeek-R1 tokenizer) of the model output.
- `correct`: label indicating if the DeepSeek-R1 generated solution matches the ground truth `solution`. Checked with [Math-Verify library](https://github.com/huggingface/Math-Verify)

## Some statistics
- The original OpenThoughts-114k dataset has **89120/113957 (78%)** math rows
- Of those, **56730/89120 (63%)** have correct answers, as checked by Math-Verify
- There is a single generation per question
- Token count distribution: mean=6366.67, std_dev=4662.88 tokens


![image/png](https://cdn-uploads.huggingface.co/production/uploads/62596f9e1c0a084224b93e00/aPYBSni3Ft6VK1VJkExtS.png)
