---
title: InternVL-U
canonical_url: "https://www.modelscope.cn/models/OpenGVLab/InternVL-U"
md_url: "https://www.modelscope.cn/models/OpenGVLab/InternVL-U.md"
repository: OpenGVLab/InternVL-U
last_updated: 2026-03-13
pipeline_tag: any-to-any
tasks:
  - any-to-any
base_model_relation: finetune
parameters: 4.3B
tensor_type:
  - BF16
  - F32
library_name:
  - safetensors
  - diffusers
  - pytorch
frameworks:
  - Pytorch
downloads: 1248
stars: 6
---

# InternVL-U

> InternVL-U - OpenGVLab 在 ModelScope 开源的模型。InternVL-U is a 4B-parameter unified multimodal model (UMM) that brings multimodal understanding, reasoning, image generation, image editing into a single framework.

- **Repository**: OpenGVLab/InternVL-U
- **Tasks**: any-to-any
- **Parameters**: 4.3B
- **Downloads**: 1248
- **Stars**: 6
- **Last updated**: 2026-03-13

Source: https://www.modelscope.cn/models/OpenGVLab/InternVL-U

---

<p align="center">
  <img src="assets/logo.jpg" width="80" />
</p>

<h1 align="center">InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing</h1>


<div align="center">

[![arXiv](https://img.shields.io/badge/ArXiv-2603.09877-b31b1b?logo=arxiv)](https://arxiv.org/abs/2603.09877)&nbsp;
[![GenEditEvalKit](https://img.shields.io/badge/GitHub-GenEditEvalKit-181717?logo=github)](https://github.com/open-compass/GenEditEvalKit)&nbsp;
[![TextEdit Benchmark](https://img.shields.io/badge/GitHub-TextEdit%20Benchmark-181717?logo=github)](https://github.com/open-compass/TextEdit)

Shanghai AI Laboratory, InternVL-U Team
</div>

**InternVL-U** is a **4B-parameter unified multimodal model (UMM)** that brings **multimodal understanding, reasoning, image generation, image editing** into a *single* framework, aiming to **democratize omni-capable multimodal intelligence** with an efficient and practical model size.

We hope **InternVL-U** serves as a **strong baseline** and accelerates progress toward **comprehensive, AGI-oriented omni-capable UMMs**.
<p align="center">
  <img src="assets/teaser1.jpg" width="40%" style="display:inline-block; vertical-align:middle;" />
  &nbsp;&nbsp;
  <img src="assets/teaser2.jpg" width="40%" style="display:inline-block; vertical-align:middle;" />
</p>


## 🤖 Model Checkpoint Download
You can download the model weights from this repository into the InternVLU project using the following command:
```code
huggingface-cli download --repo-type model --resume-download InternVL-U/InternVL-U --local-dir "your_local_path_to_store_the_model_weights"
```

## ✨Citation
If you find our InternVL-U useful, please cite our InternVL-U technical report using this BibTeX.

```
@article{tian2026internvlu,
      title={InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing},
      author={Tian, Changyao and Yang, Danni and Chen, Guanzhou and Cui, Erfei and Wang, Zhaokai and Duan, Yuchen and Yin, Penghao and Chen, Sitao and Yang, Ganlin and Liu, Mingxin and Zhu, Zirun and Fan, Ziqian and Gu, Leyao and Wang, Haomin and Wei, Qi and Yin, Jinhui and Yang, Xue and Zhong, Zhihang and Qin, Qi and Xin, Yi and Fu, Bin and Liu, Yihao and Ge, Jiaye and Guo, Qipeng and Luo, Gen and Li, Hongsheng and Qiao, Yu and Chen, Kai and Zhang, Hongjie},
      year={2026},
      eprint={2603.09877},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2603.09877}
}
```
