---
title: VoxCPM2-GGUF
canonical_url: "https://www.modelscope.cn/models/DennisHuang/VoxCPM2-GGUF"
md_url: "https://www.modelscope.cn/models/DennisHuang/VoxCPM2-GGUF.md"
repository: DennisHuang/VoxCPM2-GGUF
chinese_name: "VoxCPM2 GGUF 权重"
last_updated: 2026-06-19
license: apache-2.0
base_model:
  - openbmb/VoxCPM2
base_model_relation: quantized
library_name:
  - gguf
language:
  - zh
  - en
downloads: 757483
stars: 6
tags:
  - text-to-speech
  - tts
  - voice-cloning
  - gguf
  - llama.cpp
  - voxcpm
---

# VoxCPM2-GGUF

> VoxCPM2-GGUF - DennisHuang 在 ModelScope 开源的模型。VoxCPM2 — GGUF weights for llama.cpp-omni

DennisHuang/VoxCPM2-GGUF 是 ModelScope 魔搭社区上的机器学习模型，采用 apache-2.0 许可，基于 openbmb/VoxCPM2 构建。

- **Repository**: DennisHuang/VoxCPM2-GGUF
- **License**: apache-2.0
- **Base model**: openbmb/VoxCPM2
- **Tags**: text-to-speech, tts, voice-cloning, gguf, llama.cpp, voxcpm
- **Downloads**: 757483
- **Stars**: 6
- **Last updated**: 2026-06-19

Source: https://www.modelscope.cn/models/DennisHuang/VoxCPM2-GGUF

---

# VoxCPM2 — GGUF weights for llama.cpp-omni

GGUF-converted weights of [openbmb/VoxCPM2](https://huggingface.co/openbmb/VoxCPM2)
for the C++/ggml inference engine
[llama.cpp-omni](https://github.com/tc-mb/llama.cpp-omni) (`tools/omni/voxcpm2`).

These let you run VoxCPM2 text-to-speech and zero-shot voice cloning **natively
on CPU / Metal / CUDA / Vulkan** via ggml — no PyTorch runtime required.

## Files

| File | Format | Size | Component |
|------|--------|------|-----------|
| `VoxCPM2-BaseLM-F16.gguf` | F16 | ~3.0 GB | Base language model (28-layer, n_embd=2048) |
| `VoxCPM2-BaseLM-Q8_0.gguf` | Q8_0 | ~1.6 GB | Base language model, 8-bit quantized (recommended) |
| `VoxCPM2-Acoustic-F16.gguf` | F16 | ~1.7 GB | Acoustic stack (ResidualLM + FSQ + LocEnc/LocDiT CFM + AudioVAE) |

VoxCPM2 outputs **48 kHz mono** audio. You need **one BaseLM** (F16 or Q8_0) +
the **Acoustic** file. Q8_0 halves the BaseLM download with negligible quality
loss and is slightly faster (measured RTF 1.76 vs 1.94 on Apple M4 Pro / Metal).

## Usage

Build `voxcpm2-cli` from llama.cpp-omni, then:

```bash
# Basic TTS (GPU by default; add --cpu to force CPU)
./voxcpm2-cli \
    -t "Hello, this is VoxCPM2 running through llama.cpp-omni." \
    -o output.wav \
    VoxCPM2-BaseLM-F16.gguf \
    VoxCPM2-Acoustic-F16.gguf

# Voice cloning (reference audio)
./voxcpm2-cli -t "Cloned voice." -r speaker.wav -o clone.wav \
    VoxCPM2-BaseLM-F16.gguf VoxCPM2-Acoustic-F16.gguf

# Reference-transcript ("ultimate") cloning
./voxcpm2-cli -t "Target text." --prompt-wav speaker.wav --prompt-text "transcript of speaker.wav" \
    -o clone.wav VoxCPM2-BaseLM-F16.gguf VoxCPM2-Acoustic-F16.gguf
```

Key flags: `--cfg` (guidance scale, default 2.0), `--timesteps` (CFM steps, default 10),
`--seed`, `--temperature`, `--stream`. Voice design: prefix the text with
`(a calm female voice)…`.

## Conversion

Produced with the official converter against the upstream PyTorch weights:

```bash
python tools/omni/voxcpm2/convert_voxcpm2_to_gguf.py \
    --model model.safetensors \
    --vae audiovae.pth \
    --config config.json \
    --output ./out
```

## License & attribution

Weights derive from [openbmb/VoxCPM2](https://huggingface.co/openbmb/VoxCPM2);
their original license/terms apply. Conversion tooling and inference engine:
[llama.cpp-omni](https://github.com/tc-mb/llama.cpp-omni).
