---
title: EntroPackPreQuants
canonical_url: "https://www.modelscope.cn/models/DiffSynth-Studio/EntroPackPreQuants"
md_url: "https://www.modelscope.cn/models/DiffSynth-Studio/EntroPackPreQuants.md"
repository: DiffSynth-Studio/EntroPackPreQuants
chinese_name: "EntroPack预量化模型"
last_updated: 2026-09-29
pipeline_tag: text-to-image-synthesis
tasks:
  - text-to-image-synthesis
parameters: 271.4B
tensor_type:
  - U16
  - I64
  - U8
  - I32
  - F32
  - BF16
  - U32
library_name:
  - safetensors
downloads: 147
stars: 1
---

# EntroPackPreQuants

> EntroPackPreQuants - DiffSynth-Studio 在 ModelScope 开源的模型。EntroPack Pre-Quantized Models

- **Repository**: DiffSynth-Studio/EntroPackPreQuants
- **Tasks**: text-to-image-synthesis
- **Parameters**: 271.4B
- **Downloads**: 147
- **Stars**: 1
- **Last updated**: 2026-09-29

Source: https://www.modelscope.cn/models/DiffSynth-Studio/EntroPackPreQuants

---

# EntroPack Pre-Quantized Models

This repository releases pre-quantized weight packages produced by [**EntroPack**](https://github.com/modelscope/entropack) for representative SOTA diffusion models, to be used with the quantization features of [**DiffSynth-Studio**](https://github.com/modelscope/DiffSynth-Studio):

- **Z-Image-Turbo** (DiT + text encoder)
- **Qwen-Image-2.1** (DiT + text encoder)
- **MiniMax-H3** (two DiTs, FL2VA / Ref2VA, sharing the text encoder and video VAE)

For each model, both the **DiT and the text encoder** are provided at five fine-grained compression rates (4 / 5 / 6 / 7 / 8 bpp), plus two kinds of specialized packages:

- **`fp8@5`** (`dit_fp8_5bpp`): the weights are first **FP8 W8A8-quantized** (8-bit fp8 codes), then entropy-coded by EntroPack down to **5-bit storage** (actual bpp = 5); the forward pass uses **FP8 GEMM** for faster inference;
- **Extreme mixed-bit packages**: bit-rate is allocated per layer by sensitivity (the most sensitive layers keep a high bpp while the rest use a low bpp), pushing the whole model to a much lower bpp (Z-Image-Turbo 2.3bpp / Qwen-Image-2.1 3.0bpp / MiniMax-H3 3.0bpp).

All packages are built by entropy-coding the original weights directly with EntroPack; the metrics in the tables are weighted measurements on the package artifacts:

- **bpp** (bits per parameter): average bits per weight, `bpp = 8 × Σ storage bytes / Σ numel`.
- **Linear weight compression ratio**: derived from bpp, `ratio = bpp / 16`.
- **Weight reconstruction error** (DiT only): relative deviation between the dequantized and the original weights, `sqrt(Σ‖Wq−W‖²) / sqrt(Σ‖W‖²)`.
- **Registration entry**: each `*.safetensors` ships with a same-named `*.json`; see the end of this document for loading.

Validation outputs (images / videos) for each model can be found under `assets/`.

## Z-Image-Turbo

### DiT

| name | file | actual bpp | weight reconstruction error | size |
|---|---|---|---|---|
| dit_4bpp | `models/Z-Image-Turbo/dit_4bpp.safetensors` | 4.024 | 7.013% | 2.88 GB |
| dit_5bpp | `models/Z-Image-Turbo/dit_5bpp.safetensors` | 5.046 | 3.630% | 3.62 GB |
| dit_6bpp | `models/Z-Image-Turbo/dit_6bpp.safetensors` | 6.055 | 1.854% | 4.34 GB |
| dit_7bpp | `models/Z-Image-Turbo/dit_7bpp.safetensors` | 7.070 | 0.971% | 5.07 GB |
| dit_8bpp | `models/Z-Image-Turbo/dit_8bpp.safetensors` | 8.060 | 0.532% | 5.78 GB |
| dit_fp8_5bpp | `models/Z-Image-Turbo/dit_fp8_5bpp.safetensors` | 5.041 | 4.648% | 3.62 GB |
| dit_extreme_2.3bpp | `models/Z-Image-Turbo/dit_extreme_2.3bpp.safetensors` | 2.319 | 24.007% | 1.66 GB |
| nf4 (bitsandbytes) (for comparison) | - | 4.127 | 9.415% | - |
| torchao int8 w8a16 (for comparison) | - | 8.012 | 1.078% | - |
| torchao fp8 w8a16 (for comparison) | - | 8.010 | 2.650% | - |

### Text Encoder

| name | file | actual bpp | size |
|---|---|---|---|
| text_encoder_4bpp | `models/Z-Image-Turbo/text_encoder_4bpp.safetensors` | 4.028 | 2.43 GB |
| text_encoder_5bpp | `models/Z-Image-Turbo/text_encoder_5bpp.safetensors` | 5.051 | 2.86 GB |
| text_encoder_6bpp | `models/Z-Image-Turbo/text_encoder_6bpp.safetensors` | 6.058 | 3.29 GB |
| text_encoder_7bpp | `models/Z-Image-Turbo/text_encoder_7bpp.safetensors` | 7.061 | 3.71 GB |
| text_encoder_8bpp | `models/Z-Image-Turbo/text_encoder_8bpp.safetensors` | 8.041 | 4.13 GB |

### Extreme Mixed-Bit Package

The extreme package `dit_extreme_2.3bpp` uses mixed quantization (per-layer-group bit allocation, weighted target = 2.3bpp): modulation / embedding / adaLN / final layers (38 layers) @ **4.0bpp**, attention output projections `*.attention.to_out.0` (34 layers) @ **3.0bpp**, and the remaining 204 layers @ **2.1939bpp**. Measured 2.319bpp / weight reconstruction error 24.0%.

Example output:

![dit_extreme_2.3bpp](assets/z_image_turbo_dit_extreme_2.3bpp.jpg)


## Qwen-Image-2.1

### DiT

| name | file | actual bpp | weight reconstruction error | size |
|---|---|---|---|---|
| dit_4bpp | `models/Qwen-Image-2.1/dit_4bpp.safetensors` | 4.027 | 7.064% | 3.34 GB |
| dit_5bpp | `models/Qwen-Image-2.1/dit_5bpp.safetensors` | 5.052 | 3.658% | 4.19 GB |
| dit_6bpp | `models/Qwen-Image-2.1/dit_6bpp.safetensors` | 6.066 | 1.868% | 5.02 GB |
| dit_7bpp | `models/Qwen-Image-2.1/dit_7bpp.safetensors` | 7.085 | 0.978% | 5.87 GB |
| dit_8bpp | `models/Qwen-Image-2.1/dit_8bpp.safetensors` | 8.068 | 0.536% | 6.68 GB |
| dit_fp8_5bpp | `models/Qwen-Image-2.1/dit_fp8_5bpp.safetensors` | 5.048 | 4.662% | 4.19 GB |
| dit_extreme_3.0bpp | `models/Qwen-Image-2.1/dit_extreme_3.0bpp.safetensors` | 3.023 | 14.247% | 2.50 GB |
| nf4 (bitsandbytes) (for comparison) | - | 4.127 | 9.428% | - |
| torchao int8 w8a16 (for comparison) | - | 8.008 | 1.095% | - |
| torchao fp8 w8a16 (for comparison) | - | 8.007 | 2.648% | - |

### Text Encoder

| name | file | actual bpp | size |
|---|---|---|---|
| text_encoder_4bpp | `models/Qwen-Image-2.1/text_encoder_4bpp.safetensors` | 4.034 | 4.99 GB |
| text_encoder_5bpp | `models/Qwen-Image-2.1/text_encoder_5bpp.safetensors` | 5.062 | 5.97 GB |
| text_encoder_6bpp | `models/Qwen-Image-2.1/text_encoder_6bpp.safetensors` | 6.082 | 6.93 GB |
| text_encoder_7bpp | `models/Qwen-Image-2.1/text_encoder_7bpp.safetensors` | 7.092 | 7.89 GB |
| text_encoder_8bpp | `models/Qwen-Image-2.1/text_encoder_8bpp.safetensors` | 8.072 | 8.82 GB |

### Extreme Mixed-Bit Package

The extreme package `dit_extreme_3.0bpp` uses mixed quantization (weighted target = 3.0bpp): modulation-related layers (`modulation.1`, `norm_out.linear`, `time_text_embed.*`, 4 layers) @ **4.0bpp**, and the remaining 228 layers @ **2.9855bpp**. Measured 3.023bpp / weight reconstruction error 14.2%.

Example output:

![dit_extreme_3.0bpp](assets/qwen_image_2.1_dit_extreme_3.0bpp.png)


## MiniMax-H3

### DiT

| name | file | actual bpp | weight reconstruction error | size |
|---|---|---|---|---|
| fl2va_dit_4bpp | `models/MiniMax-H3/FL2VA/dit_4bpp.safetensors` | 4.038 | 6.567% | 15.58 GB |
| ref2va_dit_4bpp | `models/MiniMax-H3/REF2VA/dit_4bpp.safetensors` | 4.038 | 6.567% | 15.58 GB |
| fl2va_dit_5bpp | `models/MiniMax-H3/FL2VA/dit_5bpp.safetensors` | 5.065 | 3.392% | 19.54 GB |
| ref2va_dit_5bpp | `models/MiniMax-H3/REF2VA/dit_5bpp.safetensors` | 5.065 | 3.392% | 19.54 GB |
| fl2va_dit_6bpp | `models/MiniMax-H3/FL2VA/dit_6bpp.safetensors` | 6.081 | 1.734% | 23.45 GB |
| ref2va_dit_6bpp | `models/MiniMax-H3/REF2VA/dit_6bpp.safetensors` | 6.082 | 1.732% | 23.46 GB |
| fl2va_dit_7bpp | `models/MiniMax-H3/FL2VA/dit_7bpp.safetensors` | 7.093 | 0.911% | 27.36 GB |
| ref2va_dit_7bpp | `models/MiniMax-H3/REF2VA/dit_7bpp.safetensors` | 7.095 | 0.908% | 27.37 GB |
| fl2va_dit_8bpp | `models/MiniMax-H3/FL2VA/dit_8bpp.safetensors` | 8.081 | 0.500% | 31.16 GB |
| ref2va_dit_8bpp | `models/MiniMax-H3/REF2VA/dit_8bpp.safetensors` | 8.079 | 0.500% | 31.16 GB |
| fl2va_dit_fp8_5bpp | `models/MiniMax-H3/FL2VA/dit_fp8_5bpp.safetensors` | 5.058 | 4.510% | 19.54 GB |
| ref2va_dit_fp8_5bpp | `models/MiniMax-H3/REF2VA/dit_fp8_5bpp.safetensors` | 5.058 | 4.510% | 19.54 GB |
| fl2va_dit_extreme_3.0bpp | `models/MiniMax-H3/FL2VA/dit_extreme_3.0bpp.safetensors` | 3.026 | 13.225% | 11.68 GB |
| ref2va_dit_extreme_3.0bpp | `models/MiniMax-H3/REF2VA/dit_extreme_3.0bpp.safetensors` | 3.026 | 13.225% | 11.68 GB |
| nf4 (bitsandbytes) (for comparison) | - | 4.127 | 9.151% | - |
| torchao int8 w8a16 (for comparison) | - | 8.010 | 0.984% | - |
| torchao fp8 w8a16 (for comparison) | - | 8.008 | 2.642% | - |

### Text Encoder

| name | file | actual bpp | size |
|---|---|---|---|
| text_encoder_4bpp | `models/MiniMax-H3/text_encoder_4bpp.safetensors` | 4.039 | 13.20 GB |
| text_encoder_5bpp | `models/MiniMax-H3/text_encoder_5bpp.safetensors` | 5.073 | 16.21 GB |
| text_encoder_6bpp | `models/MiniMax-H3/text_encoder_6bpp.safetensors` | 6.102 | 19.20 GB |
| text_encoder_7bpp | `models/MiniMax-H3/text_encoder_7bpp.safetensors` | 7.127 | 22.18 GB |
| text_encoder_8bpp | `models/MiniMax-H3/text_encoder_8bpp.safetensors` | 8.095 | 24.99 GB |

### Video VAE

| name | file | actual bpp | size |
|---|---|---|---|
| video_vae_4bpp | `models/MiniMax-H3/video_vae_4bpp.safetensors` | 4.020 | 1.47 GB |
| video_vae_5bpp | `models/MiniMax-H3/video_vae_5bpp.safetensors` | 5.040 | 1.76 GB |
| video_vae_6bpp | `models/MiniMax-H3/video_vae_6bpp.safetensors` | 6.045 | 2.04 GB |
| video_vae_7bpp | `models/MiniMax-H3/video_vae_7bpp.safetensors` | 7.055 | 2.33 GB |
| video_vae_8bpp | `models/MiniMax-H3/video_vae_8bpp.safetensors` | 8.056 | 2.61 GB |

### Extreme Mixed-Bit Package

The extreme package `dit_extreme_3.0bpp` (identical configuration for FL2VA / Ref2VA) uses mixed quantization (weighted target = 3.0bpp): low-dimensional patch/final modules (`time_embedder.proj_in/proj_out`, `video/audio_patch_proj`, `condition_proj`, `final_layer.video_out/audio_out`, 7 layers) @ **8.0bpp**, modulation (`*.adaln_proj.linear`, 51 layers) and `attn.out_proj` (52 layers) @ **3.0bpp**, and the remaining 156 layers @ **2.9876bpp**. Measured 3.026bpp / weight reconstruction error 13.2%.

Example output:

**fl2va_dit_extreme_3.0bpp**

<video src="assets/minimax_h3_fl2va_dit_extreme_3.0bpp.mp4" controls loop></video>

**ref2va_dit_extreme_3.0bpp**

<video src="assets/minimax_h3_ref2va_dit_extreme_3.0bpp.mp4" controls loop></video>


## Usage

### Install from Source

```bash
# DiffSynth-Studio
git clone https://github.com/modelscope/DiffSynth-Studio.git
cd DiffSynth-Studio && pip install -e .

# EntroPack (pick the extra matching your CUDA version: cuda13 / cuda12)
git clone https://github.com/modelscope/entropack.git
cd entropack && pip install -e '.[cuda13]'
```

### Example Script (Z-Image-Turbo)

```python
"""
Z-Image-Turbo — EntroPack pre-quantized inference example.

Downloads the pre-quantized packages and their json sidecars from the
`DiffSynth-Studio/EntroPackPreQuants` model repo (each json is a ready-to-use
`MODEL_CONFIGS` entry), then runs inference.

Available packages (paths inside the prequant model repo):

    dit:
        models/Z-Image-Turbo/dit_4bpp.safetensors             4bpp
        models/Z-Image-Turbo/dit_5bpp.safetensors             5bpp
        models/Z-Image-Turbo/dit_6bpp.safetensors             6bpp
        models/Z-Image-Turbo/dit_7bpp.safetensors             7bpp
        models/Z-Image-Turbo/dit_8bpp.safetensors             8bpp
        models/Z-Image-Turbo/dit_extreme_2.3bpp.safetensors   mixed allocation, 2.3bpp
        models/Z-Image-Turbo/dit_fp8_5bpp.safetensors         fp8 weights, 5bpp
    text_encoder:
        models/Z-Image-Turbo/text_encoder_4bpp.safetensors    4bpp
        models/Z-Image-Turbo/text_encoder_5bpp.safetensors    5bpp
        models/Z-Image-Turbo/text_encoder_6bpp.safetensors    6bpp
        models/Z-Image-Turbo/text_encoder_7bpp.safetensors    7bpp
        models/Z-Image-Turbo/text_encoder_8bpp.safetensors    8bpp

`dit_package` and `text_encoder_package` are chosen independently: any combination works
(e.g. `dit_8bpp` with `text_encoder_4bpp`).
"""
import json

import torch

from diffsynth.configs import MODEL_CONFIGS
from diffsynth.pipelines.z_image import ZImagePipeline, ModelConfig

prequant_model_id = "DiffSynth-Studio/EntroPackPreQuants"
dit_package = "models/Z-Image-Turbo/dit_4bpp.safetensors"
text_encoder_package = "models/Z-Image-Turbo/text_encoder_4bpp.safetensors"


def register_prequantized(package_pattern):
    """Download the pre-quantized package and its json sidecar (a ready-to-use MODEL_CONFIGS entry), return a loadable ModelConfig."""
    package = ModelConfig(model_id=prequant_model_id, origin_file_pattern=package_pattern)
    package.download_if_necessary()
    sidecar = ModelConfig(model_id=prequant_model_id, origin_file_pattern=package_pattern.replace(".safetensors", ".json"))
    sidecar.download_if_necessary()
    with open(sidecar.path) as handle:
        MODEL_CONFIGS.append(json.load(handle))
    return package


pipe = ZImagePipeline.from_pretrained(
    torch_dtype=torch.bfloat16,
    device="cuda",
    model_configs=[
        register_prequantized(dit_package),
        register_prequantized(text_encoder_package),
        ModelConfig(model_id="Tongyi-MAI/Z-Image-Turbo", origin_file_pattern="vae/diffusion_pytorch_model.safetensors"),
    ],
    tokenizer_config=ModelConfig(model_id="Tongyi-MAI/Z-Image-Turbo", origin_file_pattern="tokenizer/"),
)
prompt = "Young Chinese woman in red Hanfu, intricate embroidery. Impeccable makeup, red floral forehead pattern. Elaborate high bun, golden phoenix headdress, red flowers, beads. Holds round folding fan with lady, trees, bird. Neon lightning-bolt lamp (⚡️), bright yellow glow, above extended left palm. Soft-lit outdoor night background, silhouetted tiered pagoda (西安大雁塔), blurred colorful distant lights."
image = pipe(prompt=prompt, seed=42, rand_device="cuda")
image.save("image_Z-Image-Turbo-EntroPack.jpg")
```

`dit_package` and `text_encoder_package` can be chosen independently in any combination; the extreme and fp8 packages work in exactly the same way — just point to the corresponding `.safetensors` path (mixed-quantization `configs` and other details live in the same-named `.json`, no script changes needed). For Qwen-Image-2.1 and MiniMax-H3, swap in the corresponding pipeline and package paths.
