---
title: libwaifu-cosyvoice-3
canonical_url: "https://www.modelscope.cn/models/ling0322/libwaifu-cosyvoice-3"
md_url: "https://www.modelscope.cn/models/ling0322/libwaifu-cosyvoice-3.md"
repository: ling0322/libwaifu-cosyvoice-3
last_updated: 2026-10-02
license: apache-2.0
pipeline_tag: text-to-speech
tasks:
  - text-to-speech
base_model:
  - FunAudioLLM/Fun-CosyVoice3-0.5B-2512
  - funasr/campplus
base_model_relation: finetune
parameters: 107.5M
tensor_type:
  - F32
library_name:
  - safetensors
language:
  - zh
  - en
  - fr
  - es
  - ja
  - ko
  - it
  - ru
  - de
downloads: 4
stars: 0
tags:
  - libwaifu
  - cosyvoice3
  - text-to-speech
  - voice-cloning
---

# libwaifu-cosyvoice-3

> libwaifu-cosyvoice-3 - ling0322 在 ModelScope 开源的模型。Fun-CosyVoice3-0.5B-2512 — the libwaifu package

ling0322/libwaifu-cosyvoice-3 是 ModelScope 魔搭社区上的 107.5M 参数text-to-speech模型，采用 apache-2.0 许可，基于 FunAudioLLM/Fun-CosyVoice3-0.5B-2512、funasr/campplus 构建。

- **Repository**: ling0322/libwaifu-cosyvoice-3
- **License**: apache-2.0
- **Tasks**: text-to-speech
- **Parameters**: 107.5M
- **Base model**: FunAudioLLM/Fun-CosyVoice3-0.5B-2512, funasr/campplus
- **Tags**: libwaifu, cosyvoice3, text-to-speech, voice-cloning
- **Downloads**: 4
- **Stars**: 0
- **Last updated**: 2026-10-02

Source: https://www.modelscope.cn/models/ling0322/libwaifu-cosyvoice-3

---

# Fun-CosyVoice3-0.5B-2512 — the libwaifu package

**Fun-CosyVoice3-0.5B-2512**, converted to the package format
[libwaifu](https://github.com/ling0322/libwaifu) reads. A few seconds of somebody speaking and a
sentence in; the sentence, in that voice, out, at 24 kHz. The same five models drawing the same
audio; a different file layout and one table instead of five files.

This is a **modified copy** under the original Apache License 2.0 — see `LICENSE` and `NOTICE` for
what was changed. It is **not an official FunAudioLLM product and is not endorsed by them**. The
original weights are at
[FunAudioLLM/Fun-CosyVoice3-0.5B-2512](https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512).

## What is here

| package | on disk | what it is |
|---|---|---|
| `cosyvoice3.yaml` and its parts | 4.4 GB | float32 throughout, 1 109 M parameters |

One manifest carries the whole pipeline: the Qwen2-0.5B speech-token language model, the DiT flow,
the HiFT vocoder, S3Tokenizer v3 and CAMPPlus, plus the tokenizer. Nothing else has to be fetched
to speak.

## Speaking with it

```bash
waifu -m cosyvoice                         # the page's text2speech tab
cargo run --release --example speak -- cosyvoice3.yaml voice.wav "你好。" out.wav
```

```rust
let tts = CosyVoice3::from_manifest(Device::Cuda, Residency::Device, &manifest)?;
let voice = tts.listen(&recording)?;                               // once per voice
let sound = tts.say("今天天气很好。", &voice, &options, &mut report)?;  // no transcript
let sound = tts.say_after("今天天气很好。", "录音里说的话。", &voice, &options, &mut report)?;
```

`docs/cosyvoice3.md` in the libwaifu repository has the full pipeline and how it was checked
against upstream stage by stage.

## Where this package knowingly differs from upstream

- Streaming, instruct mode (`inference_instruct2`), voice conversion and fp16 are not implemented.
- HiFT's random buffers are drawn from the reading's seed, and the flow's starting noise is
  upstream's own draw, stored as `cosyvoice3.flow.rand_noise`.

## Attribution

Fun-CosyVoice3-0.5B-2512 is licensed under the Apache License 2.0 — see
https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512. This package also bundles
**CAMPPlus** from [funasr/campplus](https://huggingface.co/funasr/campplus) (Apache License 2.0),
converted the same way. See `NOTICE` for exactly what was changed.
