---
title: OmniVoice-Onnx
canonical_url: "https://www.modelscope.cn/models/onnx-community/OmniVoice-Onnx"
md_url: "https://www.modelscope.cn/models/onnx-community/OmniVoice-Onnx.md"
repository: onnx-community/OmniVoice-Onnx
last_updated: 2026-08-01
license: apache-2.0
pipeline_tag: text-to-speech
tasks:
  - text-to-speech
model_type:
  - qwen3
architectures:
  - Qwen3ForCausalLM
base_model:
  - k2-fsa/OmniVoice
base_model_relation: finetune
library_name:
  - onnx
  - pytorch
frameworks:
  - pytorch
language:
  - aae
  - aal
  - aao
  - ab
  - abb
  - abn
  - abr
  - abs
  - abv
  - acm
  - acw
  - acx
  - adf
  - adx
  - ady
  - aeb
  - aec
  - af
  - afb
  - afo
  - ahl
  - ahs
  - ajg
  - aju
  - ala
  - aln
  - alo
  - am
  - amu
  - an
  - anc
  - ank
  - anp
  - anw
  - aom
  - apc
  - apd
  - arb
  - arq
  - ars
  - ary
  - arz
  - as
  - ast
  - avl
  - awo
  - ayl
  - ayp
  - az
  - ba
  - bag
  - bas
  - bax
  - bba
  - bbj
  - bbl
  - bbu
  - bce
  - bci
  - bcs
  - bcy
  - bda
  - bde
  - bdm
  - be
  - beb
  - bew
  - bfd
  - bft
  - bg
  - bgp
  - bhb
  - bhh
  - bho
  - bhp
  - bhr
  - bjj
  - bjk
  - bjn
  - bjt
  - bkh
  - bkm
  - bky
  - bmm
  - bmq
  - bn
  - bnm
  - bnn
  - bns
  - bo
  - bou
  - bqg
  - br
  - bra
  - brh
  - bri
  - brx
  - bs
  - bsh
  - bsj
  - bsk
  - btm
  - btv
  - bug
  - bum
  - buo
  - bux
  - bwr
  - bxf
  - byc
  - bys
  - byv
  - byx
  - bzc
  - bzw
  - ca
  - ccg
  - ceb
  - cen
  - cfa
  - cgg
  - chq
  - cjk
  - ckb
  - ckl
  - ckr
  - cky
  - cnh
  - cpy
  - cs
  - cte
  - ctl
  - cut
  - cux
  - cv
  - cy
  - da
  - dag
  - dar
  - dav
  - dbd
  - dcc
  - de
  - deg
  - dgh
  - dgo
  - dje
  - dmk
  - dml
  - dru
  - dty
  - dua
  - dv
  - dyu
  - dzg
  - ebr
  - ebu
  - ego
  - eiv
  - eko
  - ekr
  - el
  - elm
  - en
  - eo
  - es
  - esu
  - et
  - eto
  - ets
  - etu
  - eu
  - ewo
  - ext
  - eyo
  - fa
  - fan
  - fat
  - ff
  - ffm
  - fi
  - fia
  - fil
  - fip
  - fkk
  - fmp
  - fr
  - fub
  - fuc
  - fue
  - fuf
  - fuh
  - fui
  - fuq
  - fuv
  - fy
  - ga
  - gbm
  - gbr
  - gby
  - gcc
  - gdf
  - gej
  - ges
  - ggg
  - gid
  - gig
  - giz
  - gjk
  - gju
  - gl
  - glw
  - gn
  - gol
  - gom
  - gsl
  - gu
  - gui
  - gur
  - guz
  - gv
  - gwc
  - gwe
  - gwt
  - gya
  - gyz
  - ha
  - hah
  - hao
  - haw
  - haz
  - hbb
  - he
  - hem
  - hi
  - hia
  - hkk
  - hla
  - hno
  - hoj
  - hr
  - hsb
  - ht
  - hu
  - hue
  - hul
  - hux
  - hwo
  - hy
  - hz
  - ia
  - ibb
  - id
  - ida
  - idu
  - ig
  - ijc
  - ijn
  - ik
  - ikw
  - is
  - ish
  - iso
  - it
  - its
  - itw
  - itz
  - ja
  - jal
  - jax
  - jgo
  - jmx
  - jns
  - jqr
  - juk
  - juo
  - jv
  - ka
  - kab
  - kai
  - kaj
  - kam
  - kbd
  - kbl
  - kbt
  - kcq
  - kdh
  - kea
  - keu
  - kfe
  - kfk
  - kfp
  - khg
  - khw
  - kj
  - kjc
  - kjk
  - kk
  - kln
  - kls
  - km
  - kmr
  - kmy
  - kn
  - kna
  - knn
  - ko
  - kol
  - koo
  - kpo
  - kqo
  - ks
  - ksd
  - ksf
  - kto
  - kuh
  - kvx
  - kw
  - kwm
  - kxp
  - ky
  - kyx
  - lag
  - lb
  - lcm
  - ldb
  - lg
  - lij
  - lir
  - lkb
  - lla
  - ln
  - lnu
  - lo
  - loa
  - lrk
  - lss
  - lt
  - ltg
  - lto
  - lua
  - luo
  - lus
  - lv
  - lwg
  - mab
  - maf
  - mai
  - mau
  - max
  - mbo
  - mcf
  - mcn
  - mcx
  - mdd
  - mde
  - mdf
  - mek
  - mer
  - meu
  - mfm
  - mfn
  - mfo
  - mfv
  - mgg
  - mgi
  - mhk
  - mhr
  - mi
  - mig
  - miu
  - mk
  - mkf
  - mki
  - ml
  - mlq
  - mn
  - mne
  - mni
  - mqy
  - mr
  - mrj
  - mrr
  - mrt
  - ms
  - mse
  - msh
  - msw
  - mt
  - mtr
  - mtu
  - mtx
  - mua
  - mug
  - mui
  - mve
  - mvy
  - mxs
  - mxu
  - mxy
  - my
  - myv
  - mzl
  - nal
  - nan
  - nap
  - nb
  - nbh
  - ncf
  - nco
  - ncx
  - ndi
  - ng
  - ngi
  - nhg
  - nhi
  - nhn
  - nhq
  - nja
  - nl
  - nla
  - nlv
  - nmg
  - nmz
  - nn
  - nnh
  - no
  - noe
  - npi
  - nso
  - ny
  - nyu
  - oc
  - odk
  - odu
  - ogo
  - om
  - orc
  - oru
  - ory
  - os
  - pa
  - pbs
  - pbt
  - pbu
  - pcm
  - pex
  - phl
  - phr
  - pip
  - piy
  - pko
  - pl
  - plk
  - plt
  - pmq
  - pms
  - pmy
  - pnb
  - poc
  - poe
  - pow
  - prq
  - ps
  - pst
  - pt
  - pua
  - pwn
  - qug
  - qum
  - qup
  - qur
  - qus
  - quv
  - qux
  - quy
  - qva
  - qvi
  - qvj
  - qvl
  - qwa
  - qws
  - qxa
  - qxp
  - qxt
  - qxu
  - qxw
  - rag
  - rm
  - ro
  - rob
  - rof
  - roo
  - rth
  - ru
  - rup
  - rw
  - sa
  - sah
  - sat
  - sau
  - say
  - sbn
  - sc
  - scl
  - scn
  - sd
  - sei
  - shu
  - si
  - sip
  - siw
  - sjr
  - sk
  - skg
  - skr
  - sl
  - sn
  - snc
  - snk
  - so
  - sol
  - sps
  - sq
  - sr
  - src
  - sro
  - ssi
  - ste
  - sua
  - sv
  - sva
  - sw
  - szy
  - ta
  - tan
  - tar
  - tay
  - tbf
  - tcf
  - tcy
  - tdn
  - tdx
  - te
  - tg
  - tgc
  - th
  - the
  - thq
  - thr
  - thv
  - ti
  - tig
  - tio
  - tk
  - tkg
  - tkt
  - tli
  - tlp
  - tn
  - tok
  - tpl
  - tpz
  - tqp
  - tr
  - trp
  - trq
  - trv
  - trw
  - tt
  - ttj
  - ttr
  - ttu
  - tui
  - tul
  - tuq
  - tuv
  - tuy
  - tvo
  - tvu
  - tw
  - twu
  - txs
  - txy
  - udl
  - ug
  - uk
  - uki
  - umb
  - ur
  - ush
  - uz
  - uzn
  - vai
  - var
  - ver
  - vi
  - vmc
  - vmj
  - vmm
  - vmp
  - vmz
  - vot
  - vro
  - wbl
  - wci
  - weo
  - wes
  - wja
  - wji
  - wo
  - wof
  - xh
  - xhe
  - xka
  - xmf
  - xmv
  - xmw
  - xpe
  - xti
  - xtu
  - yaq
  - yav
  - yay
  - ydd
  - ydg
  - yer
  - yes
  - yi
  - yo
  - yue
  - zga
  - zgh
  - zh
  - zoc
  - zoh
  - zor
  - zpv
  - zpy
  - ztg
  - ztn
  - ztp
  - zts
  - ztu
  - zu
  - zza
downloads: 18
stars: 1
tags:
  - zero-shot
  - multilingual
  - voice-cloning
  - voice-design
  - onnx
  - onnxruntime
  - onnxruntime-genai
---

# OmniVoice-Onnx

> OmniVoice-Onnx - onnx-community 在 ModelScope 开源的模型。ONNX conversion of Prince-1/OmniVoice — a zero-shot text-to-speech model supporting 600+ languages — via Olive and onnxruntime-genai ModelBuilder.

onnx-community/OmniVoice-Onnx 是 ModelScope 魔搭社区上的text-to-speech模型，采用 apache-2.0 许可，基于 k2-fsa/OmniVoice 构建。

- **Repository**: onnx-community/OmniVoice-Onnx
- **License**: apache-2.0
- **Tasks**: text-to-speech
- **Base model**: k2-fsa/OmniVoice
- **Tags**: zero-shot, multilingual, voice-cloning, voice-design, onnx, onnxruntime, onnxruntime-genai
- **Downloads**: 18
- **Stars**: 1
- **Last updated**: 2026-08-01

Source: https://www.modelscope.cn/models/onnx-community/OmniVoice-Onnx

---

# OmniVoice → ONNX

ONNX conversion of [Prince-1/OmniVoice](https://huggingface.co/Prince-1/OmniVoice) — a zero-shot
text-to-speech model supporting 600+ languages — via [Olive](https://github.com/microsoft/Olive)
and onnxruntime-genai `ModelBuilder`.

OmniVoice is **not** a vision-language model. It is a TTS model: a Qwen3-0.6B backbone driving an
8-codebook audio codec (Higgs Audio V2 Tokenizer) through a 32-step non-autoregressive unmasking
loop. All conversion configuration is defined **in Python** (`optimize.py`) — there are no external
Olive JSON files.

## Architecture

```
Text ─▶ [audio_embeddings_encoder]  (fuses text + audio-codec token embeds; +ref codes for cloning)
              │ inputs_embeds (B,S,1024)
              ▼
        [llm_decoder]               Qwen3-0.6B, 28 layers  (exclude_embeds + exclude_lm_head)
              │ hidden_states (B,S,1024)
              ▼
        [audio_heads_decoder]       Linear 1024 → 8×1025
              │ logits (B,8,S,1025)
              ▼
        32-step iterative unmasking → audio_codes (8,T)
              │
              ▼
        [higgs_decoder]             Higgs Audio V2 Tokenizer → waveform @ 24 kHz
```

| Sub-model | In → Out | Build |
|---|---|---|
| `audio_embeddings_encoder` | ids `(B,8,S)` + mask `(B,S)` → embeds `(B,S,1024)` | Olive: PyTorch→ONNX→(int4/fp16) |
| `llm_decoder` | embeds `(B,S,1024)` → hidden `(B,S,1024)` | genai `create_model` (`exclude_embeds`+`exclude_lm_head`) |
| `audio_heads_decoder` | hidden `(B,S,1024)` → logits `(B,8,S,1025)` | Olive: PyTorch→ONNX→(int4/fp16) |
| Higgs tokenizer (×4) | waveform ↔ codec codes | Olive: PyTorch→ONNX→(fp16/fp32) |

Key config: `hidden_size=1024`, `num_codebooks=8`, `audio_vocab_size=1025` (1024 codes + mask id
1024), codebook weights `[8,8,6,6,4,4,2,2]`, 32 decoding steps, 24 kHz output.

## Why the LLM bypasses Olive

onnxruntime-genai's valid precision × execution-provider combos are **FP32/INT4 on CPU** and
**FP16/INT4 on CUDA** — there is **no FP16 LLM on CPU**. Olive's `ModelBuilder` pass therefore can't
emit a CPU fp16 LLM, so `optimize.py` calls genai `create_model` **directly** and builds the LLM
**int4 on CPU** (fp16 on GPU). The audio sub-models still go through Olive. The genai LLM declares
only `inputs_embeds` + `attention_mask` (+ KV cache) and computes positions internally — no
`position_ids` input.

## Precision profiles

| `--device` | audio sub-models | LLM | output dir |
|---|---|---|---|
| `cpu` | int4 (block-wise RTN, block 128) | int4 | `<output>` |
| `cpu_fp16` | fp16 | int4 (no fp16-CPU in genai) | `<output>` |
| `gpu` | fp16 | fp16 (CUDA EP) | `<output>` |

Higgs is exported fp16 (default) or fp32 via `--higgs-precision` — never int4 (too lossy for the DAC
codec). It is precision-shared: an int4 backbone still uses the fp16/fp32 `audio_tokenizer/`.

### Built sizes (verified)

| file | fp16 (`cpu_fp16`) | int4 (`cpu`) |
|---|---|---|
| `audio_embeddings_encoder.onnx(.data)` | 327 MB | 87 MB |
| `audio_heads_decoder.onnx` | 16.8 MB | 4.5 MB |
| `llm_decoder.onnx(.data)` | 296 MB (int4) | 296 MB (int4) |
| Higgs `audio_tokenizer/` (fp16) | ~370 MB | (shared) |

## Prerequisites

```bash
pip install -r requirements.txt          # olive-ai, onnxruntime, onnxruntime-genai, transformers 5.x
pip install omnivoice                     # registers the base OmniVoice architecture for loading
```

`omnivoice` is only needed at **build** time (to load the source checkpoint); inference needs only
`onnxruntime` + `numpy` + `soundfile`. In this repo's env the build is run isolated as
`uv run --with omnivoice python optimize.py …` to keep the base environment clean.

## Build

```bash
# CPU int4 (all sub-models int4)
python optimize.py --device cpu      --output onnx/int4 --model ./model/

# CPU fp16 (audio fp16 + LLM int4)
python optimize.py --device cpu_fp16 --output onnx      --model ./model/

# CUDA fp16 (audio fp16 + LLM fp16)   — needs an NVIDIA GPU + onnxruntime-gpu
python optimize.py --device gpu      --output onnx/cuda --model ./model/

# Higgs tokenizer only (fp16 or fp32) → <output>/audio_tokenizer
python optimize.py --higgs-only --higgs-precision fp16 --output onnx --model ./model/

# Backbone + Higgs together
python optimize.py --device cpu_fp16 --include-higgs --output onnx --model ./model/
```

`optimize.py` (1) saves OmniVoice's internal Qwen3 as a standalone `Qwen3ForCausalLM` dir, (2) builds
the two audio sub-models via inline Olive configs and the LLM via genai `create_model`, (3) copies
the tokenizer + `chat_template.jinja`, and (4) writes `omnivoice_manifest.json`. All `model_config.json`
paths are relativized to bare basenames so the folder is portable.

## Inference (CPU)

```bash
# fp16 backbone
python inference.py --model_dir onnx      --higgs_dir onnx/audio_tokenizer \
    --text "Hello from OmniVoice." --output out.wav

# int4 backbone (shares the same Higgs tokenizer — int4 dir has no audio_tokenizer/ of its own)
python inference.py --model_dir onnx/int4 --higgs_dir onnx/audio_tokenizer \
    --text "Hello from OmniVoice." --output out_int4.wav
```

`--num_audio_tokens` sets output length in frames (≈ frames × 0.04 s); `--num_steps` the unmasking
steps (the loop fills ⌈remaining/remaining_steps⌉ frames per step so every frame is decoded). Voice
cloning: add `--ref_audio ref.wav --ref_text "…"` (uses the Higgs encoders).

Standalone codec round-trip:
```bash
python higgs_inference.py --models-dir onnx/audio_tokenizer --input in.wav --output rt.wav
```

## Eval

```bash
# RTF benchmark (onnxruntime only)
python eval.py --mode rtf   --model_dir onnx      --higgs_dir onnx/audio_tokenizer
python eval.py --mode rtf   --model_dir onnx/int4 --higgs_dir onnx/audio_tokenizer

# Numerical equivalence vs PyTorch (needs `omnivoice` + the source model)
python eval.py --mode equiv --model_dir onnx      --higgs_dir onnx/audio_tokenizer
```

## Verified (CPU, this repo)

| build | inference | eval RTF |
|---|---|---|
| fp16 (`onnx/`) | ✅ WAV @ 24 kHz | 0.483 (faster than real-time) |
| int4 (`onnx/int4`) | ✅ WAV @ 24 kHz | 0.592 (faster than real-time) |
| Higgs round-trip | ✅ audio→codes(8×N)→audio | — |
| cuda (`onnx/cuda`) | not tested (no NVIDIA GPU available) | — |

int4 is slower than fp16 on CPU here (int4 weight dequant overhead without an accelerated kernel);
both are comfortably real-time.

## Notes

- **No external JSON** — every Olive config is built in Python in `optimize.py`; the old
  `cpu_and_mobile/`, `cpu_fp16/`, `cuda/`, `higgs/` `*.json` files are obsolete.
- **Tokenizer** is the standard Qwen2 fast tokenizer (`tokenizer.json`) — loaded from the local dir,
  no `trust_remote_code` / `omnivoice` needed at inference.
- **KV cache** is passed empty each step (`(B, kv_heads, 0, head_dim)`) — full-sequence forward, no
  cache reuse.
- **Higgs** uses the transformers-native `higgs_audio_v2_tokenizer` (transformers ≥ 5.4) — no
  `boson-multimodal` dependency.
```
