---
title: Nexus-Gen
canonical_url: "https://www.modelscope.cn/models/DiffSynth-Studio/Nexus-Gen"
md_url: "https://www.modelscope.cn/models/DiffSynth-Studio/Nexus-Gen.md"
repository: DiffSynth-Studio/Nexus-Gen
chinese_name: "Nexus-Gen 全模态图像生成理解"
last_updated: 2025-07-18
license: "Apache License 2.0"
pipeline_tag: any-to-any
tasks:
  - any-to-any
model_type:
  - qwen2_5_vl
architectures:
  - Qwen2_5_VLForConditionalGeneration
parameters: 8.3B
tensor_type:
  - BF16
library_name:
  - safetensors
  - pytorch
frameworks:
  - Pytorch
inference_backends:
  - "deploy_task vlm/text/emb"
  - "lmdeploy 0.9.1"
  - "sglang 0.5.2"
  - "vllm 0.9.2"
downloads: 7436
stars: 10
---

# Nexus-Gen

> Nexus-Gen - DiffSynth-Studio 在 ModelScope 开源的模型。News May 27, 2025: We fine-tuned Nexus-Gen using the BLIP-3o-60k dataset, significantly improving the model's robustness to text prompts in image generation, achieving a GenEval score of 0.79. The model…

DiffSynth-Studio/Nexus-Gen 是 ModelScope 魔搭社区上的 8.3B 参数any-to-any模型，采用 Apache License 2.0 许可，可用 deploy_task vlm/text/emb、lmdeploy 0.9.1、sglang 0.5.2 部署。

- **Repository**: DiffSynth-Studio/Nexus-Gen
- **License**: Apache License 2.0
- **Tasks**: any-to-any
- **Parameters**: 8.3B
- **Inference backends**: deploy_task vlm/text/emb, lmdeploy 0.9.1, sglang 0.5.2, vllm 0.9.2
- **Downloads**: 7436
- **Stars**: 10
- **Last updated**: 2025-07-18

Source: https://www.modelscope.cn/models/DiffSynth-Studio/Nexus-Gen

---

## News
- **May 27, 2025**: We fine-tuned Nexus-Gen using the [BLIP-3o-60k](https://www.modelscope.cn/datasets/BLIP3o/BLIP3o-60k) dataset, significantly improving the model's robustness to text prompts in image generation, **achieving a GenEval score of 0.79**. The [model checkpoints](https://www.modelscope.cn/models/DiffSynth-Studio/Nexus-Gen) have been updated.

## What is the Nexus-Gen
Nexus-Gen is a unified model that synergizes the language reasoning capabilities of LLMs with the image synthesis power of diffusion models. To align the embedding space of the LLM and diffusion model, we conduct a dual-phase alignment training process. (1) The autoregressive LLM learns to predict image embeddings conditioned on multimodal inputs, while (2) the vision decoder is trained to reconstruct high-fidelity images from these embeddings. During training the LLM, we identified a critical discrepancy between the autoregressive paradigm's training and inference phases, where error accumulation in continuous embedding space severely degrades generation quality. To avoid this issue, we introduce a prefilled autoregression strategy that prefills input sequence with position-embedded special tokens instead of continuous embeddings. Through dual-phase training, Nexus-Gen has developed the integrated capability to comprehensively address the image understanding, generation and editing tasks as follows.

Online Demo: https://www.modelscope.cn/studios/DiffSynth-Studio/Nexus-Gen

More information please refer to our repo: https://github.com/modelscope/Nexus-Gen.git

![cover](assets/illustrations/gen_edit.jpg)
![architecture](assets/illustrations/architecture.png)

## Getting Started
### Installation
1. Install [DiffSynth-Studio](https://github.com/modelscope/DiffSynth-Studio.git) from source:
```shell
git clone https://github.com/modelscope/DiffSynth-Studio.git
cd DiffSynth-Studio
pip install -e .
```
2. Install requirements
```
pip install -r requirements.txt
```
3. Install [ms-swift](https://github.com/modelscope/ms-swift.git) if you want to perform finetuning on Nexus-Gen.
```
pip install ms-swift -U
```
### Prepare models
```shell
python download_models.py
```
### Image Understanding
```shell
python image_understanding.py
```

### Image Generation
image generation with detailed prompt.
```shell
python image_generation.py
```
Polish prompt and generate images with Nexus-Gen.
```shell
image_generation_with_selfpolish.py
```

### Image Editing
```shell
python image_editing.py
```

### Training Codes
Nexus-Gen is trained base on [ms-swift](https://github.com/modelscope/ms-swift.git) and [DiffSynth-Studio](https://github.com/modelscope/DiffSynth-Studio.git). You can find the training scripts in `train/scripts/train_decoder.sh` and `train_llm.sh`.


## Limitations

- Please note that Nexus-Gen is trained primarily with English corpus, therefore instruction-following with non-English is not supported.
- Please note that Nexus-Gen was trained on limited text-to-image data and may not be robust to short prompts.

## Citation

Feel free to reference our work if you find it helpful.

```
@misc{zhang2025nexusgenunifiedmodelimage,
      title={Nexus-Gen: A Unified Model for Image Understanding, Generation, and Editing}, 
      author={Hong Zhang and Zhongjie Duan and Xingjun Wang and Yuze Zhao and Weiyi Lu and Zhipeng Di and Yixuan Xu and Yingda Chen and Yu Zhang},
      year={2025},
      eprint={2504.21356},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2504.21356v2}, 
}
```
