---
title: Step-Audio-Tokenizer
canonical_url: "https://www.modelscope.cn/models/stepfun-ai/Step-Audio-Tokenizer"
md_url: "https://www.modelscope.cn/models/stepfun-ai/Step-Audio-Tokenizer.md"
repository: stepfun-ai/Step-Audio-Tokenizer
last_updated: 2025-02-18
license: apache-2.0
pipeline_tag: text-to-speech
tasks:
  - text-to-speech
library_name:
  - onnx
  - pytorch
frameworks:
  - pytorch
downloads: 6149
stars: 12
---

# Step-Audio-Tokenizer

> Step-Audio-Tokenizer - stepfun-ai 在 ModelScope 开源的模型。Step-Audio-Tokenizer

stepfun-ai/Step-Audio-Tokenizer 是 ModelScope 魔搭社区上的text-to-speech模型，采用 apache-2.0 许可。

- **Repository**: stepfun-ai/Step-Audio-Tokenizer
- **License**: apache-2.0
- **Tasks**: text-to-speech
- **Downloads**: 6149
- **Stars**: 12
- **Last updated**: 2025-02-18

Source: https://www.modelscope.cn/models/stepfun-ai/Step-Audio-Tokenizer

---

# Step-Audio-Tokenizer


Step-Audio LLM is the industry’s first 130-billion parameter hu-manlike unified end-to-end model that integrates multimodal speech un-derstanding and generation capabilities, including singing voice synthesis, tool utilization, role-play and multilingual/dialectal comprehension and synthesis. 

This repository provides the speech tokenizer component of Step-Audio LLM. For linguistic tokenization, we utilize the output from the Paraformer encoder, which is quantized into discrete representations at a token rate of 16.7 Hz. For semantic tokenization, we employ CosyVoice’s tokenizer, specifically designed to efficiently encode features essential for generating natural and expressive speech outputs, operating at a token rate of 25 Hz.

## More information
For more information, please refer to our repository: [Step-Audio](https://github.com/stepfun-ai/Step-Audio).
