---
title: fineweb-edu-1B
canonical_url: "https://www.modelscope.cn/datasets/codelion/fineweb-edu-1B"
md_url: "https://www.modelscope.cn/datasets/codelion/fineweb-edu-1B.md"
repository: codelion/fineweb-edu-1B
last_updated: 2025-11-03
license: "Apache License 2.0"
storage_size: "2.6 GB"
downloads: 193
stars: 0
---

# fineweb-edu-1B

> fineweb-edu-1B - codelion 在 ModelScope 开源的数据集。Sampling Methodology

codelion/fineweb-edu-1B 是 ModelScope 魔搭社区上的数据集，存储大小 2.6 GB，采用 Apache License 2.0 许可。

- **Repository**: codelion/fineweb-edu-1B
- **License**: Apache License 2.0
- **Storage size**: 2.6 GB
- **Downloads**: 193
- **Stars**: 0
- **Last updated**: 2025-11-03

Source: https://www.modelscope.cn/datasets/codelion/fineweb-edu-1B

---

## Sampling Methodology

This dataset was created using **reservoir sampling**, a statistically unbiased random sampling algorithm that guarantees each sample from the source dataset has an equal probability of being included. This ensures the 1B token sample is representative of the full dataset's characteristics.

**Source Dataset**: [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu)
**Sample Size**: 1B tokens
**Content**: Curated educational web resources

Reservoir sampling enables rapid experimentation and ablation studies without processing the entire source dataset, while maintaining statistical validity of results.

For details on how this dataset was used in optimal pre-training data composition research, see the [blog post](https://huggingface.co/blog/codelion/optimal-dataset-mixing/).

## Citation

If you use this model/dataset, please cite:

```bibtex
@article{sharma2025billion,
  title={The 1 Billion Token Challenge: Finding the Perfect Pre-training Mix},
  author={Sharma, Asankhaya},
  year={2025},
  url={https://huggingface.co/blog/codelion/optimal-dataset-mixing/}
}
```

For more details, see the [blog post](https://huggingface.co/blog/codelion/optimal-dataset-mixing/).
