---
title: SmallThoughts
canonical_url: "https://www.modelscope.cn/datasets/AI-ModelScope/SmallThoughts"
md_url: "https://www.modelscope.cn/datasets/AI-ModelScope/SmallThoughts.md"
repository: AI-ModelScope/SmallThoughts
last_updated: 2025-07-24
license: apache-2.0
storage_size: "175 MB"
downloads: 1097
stars: 2
---

# SmallThoughts

> SmallThoughts - AI-ModelScope 在 ModelScope 开源的数据集。Open synthetic reasoning dataset, covering math, science, code, and puzzles.

AI-ModelScope/SmallThoughts 是 ModelScope 魔搭社区上的数据集，存储大小 175 MB，采用 apache-2.0 许可。

- **Repository**: AI-ModelScope/SmallThoughts
- **License**: apache-2.0
- **Storage size**: 175 MB
- **Downloads**: 1097
- **Stars**: 2
- **Last updated**: 2025-07-24

Source: https://www.modelscope.cn/datasets/AI-ModelScope/SmallThoughts

---

# SmallThoughts

<div align="center">
  <img src="https://huggingface.co/datasets/SmallDoge/SmallThoughts/resolve/main/SmallThoughts.png" alt="Small-Thoughts Map" width="60%"/>
</div>
<div align="center">
  <a href="https://discord.gg/P2yYH95N" target="_blank" style="margin: 2px;">
    <img alt="Discord" src="https://img.shields.io/badge/Discord-Small%20Doges-7289da?logo=discord&logoColor=white&color=7289da" style="display: inline-block; vertical-align: middle;"/>
  </a>
  <a href="https://github.com/SmallDoges/small-thoughts" target="_blank" style="margin: 2px;">
    <img alt="GitHub" src="https://img.shields.io/badge/GitHub-SmallThoughts-181717?logo=github" style="display: inline-block; vertical-align: middle;"/>
  </a>
  <a href="https://github.com/SmallDoges/small-doge/blob/main/LICENSE" style="margin: 2px;">
    <img alt="License" src="https://img.shields.io/badge/License-Apache--2.0-blue.svg" style="display: inline-block; vertical-align: middle;"/>
  </a>
</div>

---

Open synthetic reasoning dataset, covering math, science, code, and puzzles.

To address the issue of the existing DeepSeek R1 distilled data being too long, this dataset constrains the reasoning trajectory to be more precise and concise while retaining the reflective nature.

We also open-sourced the pipeline code for distilled data [here](https://github.com/SmallDoges/small-thoughts), with just one command you can generate your own dataset.


## How to use

You can load the dataset with the following code:

```python
import datasets
dataset = datasets.load_dataset("SmallDoge/SmallThoughts")
```

If you are using [TRL](https://github.com/huggingface/trl) for model training, The `problem` and `solution` columns can be used for **GRPO** reinforcement learning, and the `messages` columns can be used for **SFT** fine-tuning, without any additional preprocessing.


## Visualization


All examples, clustered by semantic similarity, can be explored in [Nomic Atlas](https://atlas.nomic.ai/data/losercheems/smallthoughts/map). 

<a href="https://atlas.nomic.ai/data/losercheems/smallthoughts/map">
  <img src="https://huggingface.co/datasets/SmallDoge/SmallThoughts/resolve/main/small_thoughts_map.png" alt="Nomic Atlas Small-Thoughts Map" width="40%"/>
</a>


# License

This dataset is released under the Apache-2.0 License.


# Citation

```bibtex
@misc{wu2025concisereasoningbiggains,
      title={Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting}, 
      author={Yifan Wu and Jingze Shi and Bingheng Wu and Jiayi Zhang and Xiaotian Lin and Nan Tang and Yuyu Luo},
      year={2025},
      eprint={2505.19716},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2505.19716}, 
}
```
