---
title: StepEval-Audio-360
canonical_url: "https://www.modelscope.cn/datasets/stepfun-ai/StepEval-Audio-360"
md_url: "https://www.modelscope.cn/datasets/stepfun-ai/StepEval-Audio-360.md"
repository: stepfun-ai/StepEval-Audio-360
last_updated: 2025-04-25
license: "Apache License 2.0"
storage_size: "159 MB"
domain:
  - audio
tasks:
  - automatic_speech_recognition
languages:
  - cn
  - en
  - jp
downloads: 924
stars: 1
---

# StepEval-Audio-360

> StepEval-Audio-360 - stepfun-ai 在 ModelScope 开源的数据集。StepEval-Audio-360 Dataset Description StepEval Audio 360 is a comprehensive dataset that evaluates the ability of multi-modal large language models (MLLMs) in human-AI audio interaction. This audio…

stepfun-ai/StepEval-Audio-360 是 ModelScope 魔搭社区上的automatic_speech_recognition数据集，涉及 audio 领域，存储大小 159 MB，采用 Apache License 2.0 许可。

- **Repository**: stepfun-ai/StepEval-Audio-360
- **License**: Apache License 2.0
- **Tasks**: automatic_speech_recognition
- **Domain**: audio
- **Storage size**: 159 MB
- **Downloads**: 924
- **Stars**: 1
- **Last updated**: 2025-04-25

Source: https://www.modelscope.cn/datasets/stepfun-ai/StepEval-Audio-360

---

# StepEval-Audio-360
## Dataset Description
StepEval Audio 360 is a comprehensive dataset that evaluates the ability of multi-modal large language models (MLLMs) in human-AI audio interaction. This audio benchmark dataset, sourced from professional human annotators, covers a full spectrum of capabilities: singing, creativity, role-playing, logical reasoning, voice understanding, voice instruction following, gaming, speech emotion control, and language ability.

## Languages
StepEval Audio 360 comprises about human voice recorded in different languages and dialects, including Chinese(Szechuan dialect and cantonese), English, and Japanese. It contains both audio and transcription data.

## Links
- Homepage: [Step-Audio](https://github.com/stepfun-ai/Step-Audio)
- Paper: [Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
](https://arxiv.org/abs/2502.11946)
- ModelScope: https://modelscope.cn/datasets/stepfun-ai/StepEval-Audio-360
- Step-Audio Model Suite：
  - Step-Audio-Tokenizer：
    - Hugging Face：https://huggingface.co/stepfun-ai/Step-Audio-Tokenizer
    - ModelScope：https://modelscope.cn/models/stepfun-ai/Step-Audio-Tokenizer
  - Step-Audio-Chat :
    - HuggingFace: https://huggingface.co/stepfun-ai/Step-Audio-Chat
    - ModelScope: https://modelscope.cn/models/stepfun-ai/Step-Audio-Chat
  - Step-Audio-TTS-3B：
    - Hugging Face: https://huggingface.co/stepfun-ai/Step-Audio-TTS-3B
    - ModelScope: https://modelscope.cn/models/stepfun-ai/Step-Audio-TTS-3B

## User Manual
* Download the dataset
```
# Make sure you have git-lfs installed (https://git-lfs.com)
git lfs install
git clone https://huggingface.co/datasets/stepfun-ai/StepEval-Audio-360
cd StepEval-Audio-360
git lfs pull
```

* Decompress audio data
```
mkdir audios
tar -xvf audios.tar.gz -C audios
```

* How to use
```
from datasets import load_dataset

dataset = load_dataset("stepfun-ai/StepEval-Audio-360")
dataset = dataset["test"]
for item in dataset:
    conversation_id = item["conversation_id"]
    category = item["category"]
    conversation = item["conversation"]

    # parse multi-turn dialogue data
    for turn in conversation:
        role = turn["role"]
        text = turn["text"]
        audio_filename = turn["audio_filename"] # refer to decompressed audio file
        if audio_filename is not None:
            print(role, text, audio_filename)
        else:
            print(role, text)
```

## Licensing
This dataset project is licensed under the [Apache 2.0 License](https://www.apache.org/licenses/LICENSE-2.0).

## Citation
If you utilize this dataset, please cite it using the BibTeX provided.
```
@misc {stepfun_2025,
        author       = { {StepFun} },
        title        = { StepEval-Audio-360 (Revision 72a072e) },
        year         = 2025,
        url          = { https://huggingface.co/datasets/stepfun-ai/StepEval-Audio-360 },
        doi          = { 10.57967/hf/4528 },
        publisher    = { Hugging Face }
}
```
