---
title: DiaMoE-TTS_IPA_Trainingset
canonical_url: "https://www.modelscope.cn/datasets/giantailab/DiaMoE-TTS_IPA_Trainingset"
md_url: "https://www.modelscope.cn/datasets/giantailab/DiaMoE-TTS_IPA_Trainingset.md"
repository: giantailab/DiaMoE-TTS_IPA_Trainingset
chinese_name: DiaMoE-TTS_IPA_Trainingset
last_updated: 2025-11-28
license: cc-by-4.0
storage_size: "314 MB"
downloads: 76
stars: 0
---

# DiaMoE-TTS_IPA_Trainingset

> DiaMoE-TTS_IPA_Trainingset - giantailab 在 ModelScope 开源的数据集。DiaMoE-TTS: A Unified IPA-based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation

giantailab/DiaMoE-TTS_IPA_Trainingset 是 ModelScope 魔搭社区上的数据集，存储大小 314 MB，采用 cc-by-4.0 许可。

- **Repository**: giantailab/DiaMoE-TTS_IPA_Trainingset
- **License**: cc-by-4.0
- **Storage size**: 314 MB
- **Downloads**: 76
- **Stars**: 0
- **Last updated**: 2025-11-28

Source: https://www.modelscope.cn/datasets/giantailab/DiaMoE-TTS_IPA_Trainingset

---

# DiaMoE-TTS: A Unified IPA-based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation

github:[DiaMoE-TTS](https://github.com/GiantAILab/DiaMoE-TTS)

We utilize the [Common Voice Cantonese dataset](https://arxiv.org/abs/1912.06670), the [Emilia Mandarin dataset](https://arxiv.org/abs/2407.05361) and dialectal data from the [KeSpeech corpus](https://openreview.net/forum?id=b3Zoeq2sCLq) and a open-source [Sourthern Min dataset](https://sutian.moe.edu.tw/zh-hant/siongkuantsuguan/) for training. 
We only release the frontend of the open-source dataset IPA here.

## Short Intro

Dialect speech embodies rich cultural and linguistic diversity, yet building text-to-speech (TTS) systems for dialects remains challenging due to scarce data, inconsistent orthographies, and complex phonetic variation. To address these issues, we present DiaMoE-TTS, a unified IPA-based framework that standardizes phonetic representations and resolves grapheme-to-phoneme ambiguities. Built upon the F5-TTS architecture, the system introduces a dialect-aware Mixture-of-Experts (MoE) to model phonological differences and employs parameter-efficient adaptation with Low-Rank Adaptors (LoRA) and Conditioning Adapters for rapid transfer to new dialects. Unlike approaches dependent on large-scale or proprietary resources, DiaMoE-TTS enables scalable, open-data-driven synthesis. Experiments demonstrate natural and expressive speech generation, achieving zero-shot performance on unseen dialects and specialized domains such as Peking Opera with only a few hours of data.
