---
title: s1-m_beta
canonical_url: "https://www.modelscope.cn/datasets/PKU-Alignment/s1-m_beta"
md_url: "https://www.modelscope.cn/datasets/PKU-Alignment/s1-m_beta.md"
repository: PKU-Alignment/s1-m_beta
last_updated: 2025-03-15
license: "Apache License 2.0"
storage_size: "5.6 GB"
downloads: 1190
stars: 0
---

# s1-m_beta

> s1-m_beta - PKU-Alignment 在 ModelScope 开源的数据集。🏠 Homepage | 👍 Our Official Code Repo | 🤗 S1-M-7B Model (Beta)

PKU-Alignment/s1-m_beta 是 ModelScope 魔搭社区上的数据集，存储大小 5.6 GB，采用 Apache License 2.0 许可。

- **Repository**: PKU-Alignment/s1-m_beta
- **License**: Apache License 2.0
- **Storage size**: 5.6 GB
- **Downloads**: 1190
- **Stars**: 0
- **Last updated**: 2025-03-15

Source: https://www.modelscope.cn/datasets/PKU-Alignment/s1-m_beta

---

# S1-M Dataset (Beta)

[🏠 Homepage](https://github.com/PKU-Alignment/s1-m) | [👍 Our Official Code Repo](https://github.com/PKU-Alignment/s1-m) | [🤗 S1-M-7B Model (Beta)](https://huggingface.co/PKU-Alignment/s1-m_7b_beta) 

S1-M Dataset (Beta) is an open-source TI2T reasoning dataset used to train the S1-M Model (Beta), giving it a "think first, then response" paradigm. The prompts and images in the S1-M Dataset (Beta) come from two open-source datasets: [align-anything](https://huggingface.co/datasets/PKU-Alignment/align-anything) and [multimodal-open-r1-8k-verified](https://huggingface.co/datasets/lmms-lab/multimodal-open-r1-8k-verified), accounting for 49.62% and 50.38% respectively, aiming to balance the model's general capabilities with mathematical abilities. Data annotation uses **Claude 3.7 Sonnet 20250219** as the annotation model, which is guided to think first and then provide answers through a system prompt as shown below.

```
You are a reasoning model with advanced analytical capabilities. I will provide an image and ask a question about it. Your task is to analyze the image thoroughly and answer my question accurately.

Response format:
<think>
[step-by-step reasoning process]
</think>
[final answer]

Guidelines:

1. Place your reasoning process between <think> and </think> tags first, and the private answer after that.

2. The reasoning process can include expressions like "let me think," "oh, I see,", "maybe I should think about it from a different angle," or other natural language thought expressions.

3. For multiple-choice questions, end with "Answer: [LETTER]" where LETTER corresponds to your selected option.

Remember to be thorough in your analysis but concise in your final answer.
```

The system prompt requires Claude 3.7 to first place its thinking process between the thinking markers `<think>` and `</think>`, and then provide the final answer based on this thinking, forming a "think first, then response" paradigm.

Through this thinking process, the annotated responses have a longer token distribution. The length distribution of thinking content + answer content in the S1-M Dataset (Beta) is shown in the figure below.
![Token Length Distribution](token_length_distribution.png)

**Note: The S1-M Dataset (Beta) is still under development and the final version has not yet been released.**
