---
title: Evol-Instruct-Python-1k
canonical_url: "https://www.modelscope.cn/datasets/mlabonne/Evol-Instruct-Python-1k"
md_url: "https://www.modelscope.cn/datasets/mlabonne/Evol-Instruct-Python-1k.md"
repository: mlabonne/Evol-Instruct-Python-1k
last_updated: 2025-03-18
license: "Apache License 2.0"
storage_size: "2.2 MB"
downloads: 1008
stars: 0
---

# Evol-Instruct-Python-1k

> Evol-Instruct-Python-1k - mlabonne 在 ModelScope 开源的数据集。Evol-Instruct-Python-1k

mlabonne/Evol-Instruct-Python-1k 是 ModelScope 魔搭社区上的数据集，存储大小 2.2 MB，采用 Apache License 2.0 许可。

- **Repository**: mlabonne/Evol-Instruct-Python-1k
- **License**: Apache License 2.0
- **Storage size**: 2.2 MB
- **Downloads**: 1008
- **Stars**: 0
- **Last updated**: 2025-03-18

Source: https://www.modelscope.cn/datasets/mlabonne/Evol-Instruct-Python-1k

---

# Evol-Instruct-Python-1k

Subset of the [`mlabonne/Evol-Instruct-Python-26k`](https://huggingface.co/datasets/mlabonne/Evol-Instruct-Python-26k) dataset with only 1000 samples.

It was made by filtering out a few rows (instruction + output) with more than 2048 tokens, and then by keeping the 1000 longest samples.

Here is the distribution of the number of tokens in each row using Llama's tokenizer:

![](https://i.imgur.com/nwJbg7S.png)
