---
title: z-image-turbo-flow-dpo
canonical_url: "https://www.modelscope.cn/models/FFFFFFoo/z-image-turbo-flow-dpo"
md_url: "https://www.modelscope.cn/models/FFFFFFoo/z-image-turbo-flow-dpo.md"
repository: FFFFFFoo/z-image-turbo-flow-dpo
last_updated: 2026-02-28
license: apache-2.0
pipeline_tag: feature-extraction
tasks:
  - feature-extraction
base_model:
  - Tongyi-MAI/Z-Image-Turbo
base_model_relation: adapter
parameters: 84.8M
tensor_type:
  - BF16
library_name:
  - lora
  - safetensors
  - pytorch
frameworks:
  - Pytorch
supports_inference: txt2img
downloads: 1628
stars: 63
tags:
  - text-to-image
  - flow-matching
  - diffusion
  - lora
  - dpo
  - flow-dpo
  - z-image
  - photorealistic
  - lighting
---

# z-image-turbo-flow-dpo

> z-image-turbo-flow-dpo - FFFFFFoo 在 ModelScope 开源的模型。这是一个专门为 Z-Image-Turbo 设计的 LoRA ，使用 Flow-DPO（用于 Flow Matching 的直接偏好优化）进行微调，显著提升了照片级真实感光照、电影级阴影以及整体图像质量。通过在完美空间对齐的图像对上使用 Flow-DPO，该 LoRA 修复了步数蒸馏模型中常见的"扁平"、"褪色"或"塑料"伪影，仅需 8 次推理步骤 即可呈现令人惊叹、物理上精确的光照效果。

FFFFFFoo/z-image-turbo-flow-dpo 是 ModelScope 魔搭社区上的 84.8M 参数feature-extraction模型，采用 apache-2.0 许可，基于 Tongyi-MAI/Z-Image-Turbo 构建，并支持在线推理（txt2img）。

- **Repository**: FFFFFFoo/z-image-turbo-flow-dpo
- **License**: apache-2.0
- **Tasks**: feature-extraction
- **Parameters**: 84.8M
- **Base model**: Tongyi-MAI/Z-Image-Turbo
- **Online inference**: txt2img
- **Tags**: text-to-image, flow-matching, diffusion, lora, dpo, flow-dpo, z-image, photorealistic, lighting
- **Downloads**: 1628
- **Stars**: 63
- **Last updated**: 2026-02-28

Source: https://www.modelscope.cn/models/FFFFFFoo/z-image-turbo-flow-dpo

---

# Z-Image-Turbo Photorealistic Lighting LoRA (Flow-DPO)

This is a specialized LoRA adapter for [Alibaba-Tongyi/Z-Image-Turbo](https://huggingface.co/Alibaba-Tongyi/Z-Image-Turbo), finetuned using **Flow-DPO** (Direct Preference Optimization for Flow Matching) to significantly enhance photorealistic lighting, cinematic shadows, and overall image quality.

By utilizing Flow-DPO on perfectly spatially-aligned image pairs, this LoRA fixes the common "flat," "washed-out," or "plastic" artifacts often found in ultra-fast distilled models, delivering stunning, physically accurate lighting in just **8 inference steps**.


## 🧠 Training Details & Methodology

This model was trained using a custom implementation of **Flow-DPO** ([Improving Video Generation with Human Feedback, arXiv:2501.13918](https://arxiv.org/abs/2501.13918)).

### 1. The Dataset (Strict Spatial Alignment)
To prevent the model from hallucinating or altering image structures (Catastrophic Forgetting), the preference dataset was constructed using strict spatial alignment:
* **Win (Chosen):** High-quality, professional photographs with perfect lighting and textures.
* **Lose (Rejected):** The exact same images degraded programmatically (Gaussian blur, lowered contrast, extreme exposure shifts, gaussian noise, and heavy JPEG compression artifacts).
* **Alignment:** No cropping or warping was applied, ensuring the Flow Matching trajectory learned to solely correct lighting and texture.

### 2. Discrete Timestep Distillation Preservation
Unlike standard diffusion models where $t$ is sampled continuously $t \in [0, 1]$, Z-Image-Turbo is a **distilled model** specifically optimized for 8 fixed timesteps. 
During the Flow-DPO training, we dynamically extracted the exact discrete $t$-distribution from the `FlowMatchEulerDiscreteScheduler` and restricted the random sampling to these exact 8 nodes. This ensures the LoRA retains the turbo model's extreme speed without causing output blurriness.

### 3. Hyperparameters
* **Base Model:** Alibaba-Tongyi/Z-Image-Turbo (6B Single-Stream DiT)
* **Learning Rate:** `1e-4`
* **KL Penalty ($\beta$):** `1.0`
* **Effective Batch Size:** `1`
* **Mixed Precision:** `bfloat16`

## ⚠️ Limitations
* **Not an Image-to-Image Restorer:** This LoRA changes the *prior distribution* of the Text-to-Image generation. It is designed to generate better original images from text prompts, not to be used as an img2img filter to fix user-uploaded bad photos (unless combined with RF-Inversion techniques, which are highly unstable for 8-step models).
* **Color Saturation:** Pushing the LoRA scale too high (e.g., > 1.5) might result in over-sharpened or overly saturated images due to the nature of DPO margin maximization. Keep the scale around `0.6 - 1.0` for the most photorealistic results.

## 📚 Citation
If you find this model or training methodology useful, please consider referencing:

```bibtex
@article{liu2025improving,
  title={Improving video generation with human feedback},
  author={Liu, Jie and Liu, Gongye and Liang, Jiajun and Yuan, Ziyang and Liu, Xiaokun and Zheng, Mingwu and Wu, Xiele and Wang, Qiulin and Qin, Wenyu and Xia, Menghan and others},
  journal={arXiv preprint arXiv:2501.13918},
  year={2025}
}
