---
title: SmolLM2-135M-10B
canonical_url: "https://www.modelscope.cn/datasets/EleutherAI/SmolLM2-135M-10B"
md_url: "https://www.modelscope.cn/datasets/EleutherAI/SmolLM2-135M-10B.md"
repository: EleutherAI/SmolLM2-135M-10B
last_updated: 2025-08-15
license: "Apache License 2.0"
storage_size: "24 GB"
downloads: 160
stars: 0
---

# SmolLM2-135M-10B

> SmolLM2-135M-10B - EleutherAI 在 ModelScope 开源的数据集。This dataset is sampled from the SmolLM2 Corpus described in https://arxiv.org/abs/2502.02737. Specifically, we sampled from the SmolLM2-135M pretraining data, a 2T token mixture consisting of four complete…

EleutherAI/SmolLM2-135M-10B 是 ModelScope 魔搭社区上的数据集，存储大小 24 GB，采用 Apache License 2.0 许可。

- **Repository**: EleutherAI/SmolLM2-135M-10B
- **License**: Apache License 2.0
- **Storage size**: 24 GB
- **Downloads**: 160
- **Stars**: 0
- **Last updated**: 2025-08-15

Source: https://www.modelscope.cn/datasets/EleutherAI/SmolLM2-135M-10B

---

This dataset is sampled from the SmolLM2 Corpus described in https://arxiv.org/abs/2502.02737. Specifically, we sampled from
the SmolLM2-135M pretraining data, a 2T token mixture consisting of four complete high quality datasets, and selected portions of 
DCLM-Edu and FineWeb-Edu sampled at a 6:4 ratio.

This sample is intended to enable fast downloading and training of [sparsify](https://github.com/EleutherAI/sparsify) models.

- FineMath: 34B tokens
- Stack-Edu: 125B tokens
- InfiMM-WebMath: 40B tokens
- Cosmopedia V2: 30B tokens
- FineWeb-Edu: 710.4B tokens (1.2T in full dataset)
- DCLM-Edu: 1065.6B tokens (3.8T in full dataset)

This sample does not include the following datasets used in the otherwise similar Stage 4 of SmolLM2-1.7B training:
- [OpenWebMath](https://huggingface.co/datasets/open-web-math/open-web-math): 12B tokens
- [AugGSM8K](https://github.com/OFA-Sys/gsm8k-ScRel/tree/main/data/MuggleMATH): ?
