---
title: YingMusic-Singer
canonical_url: "https://www.modelscope.cn/models/giantailab/YingMusic-Singer"
md_url: "https://www.modelscope.cn/models/giantailab/YingMusic-Singer.md"
repository: giantailab/YingMusic-Singer
chinese_name: YingMusic-Singer
last_updated: 2026-02-10
license: cc-by-nc-4.0
pipeline_tag: text-to-speech
tasks:
  - text-to-speech
library_name:
  - pytorch
frameworks:
  - Pytorch
downloads: 30
stars: 0
---

# YingMusic-Singer

> YingMusic-Singer - giantailab 在 ModelScope 开源的模型。YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance

giantailab/YingMusic-Singer 是 ModelScope 魔搭社区上的text-to-speech模型，采用 cc-by-nc-4.0 许可。

- **Repository**: giantailab/YingMusic-Singer
- **License**: cc-by-nc-4.0
- **Tasks**: text-to-speech
- **Downloads**: 30
- **Stars**: 0
- **Last updated**: 2026-02-10

Source: https://www.modelscope.cn/models/giantailab/YingMusic-Singer

---

# YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance

github:[YingMusic-Singer](https://github.com/GiantAILab/YingMusic-Singer)

## Short Intro

YingMusic-Singer is a unified framework for Zero-shot Singing Voice Synthesis (SVS) and Editing, driven by Annotation-free Melody Guidance. Addressing the scalability challenges of real-world applications, our system eliminates the reliance on costly phoneme-level alignment and manual melody annotations. It enables arbitrary lyrics to be synthesized or edited with any reference melody in a zero-shot manner.
Our approach leverages a Diffusion Transformer (DiT) based generative model, incorporating a pre-trained melody extraction module to derive MIDI information directly from reference audio. By introducing a structured guidance mechanism and employing Flow-GRPO reinforcement learning, we achieve superior pronunciation clarity, melodic accuracy, and musicality without requiring fine-grained alignment.
