---
title: GUI-Owl-7B
canonical_url: "https://www.modelscope.cn/models/iic/GUI-Owl-7B"
md_url: "https://www.modelscope.cn/models/iic/GUI-Owl-7B.md"
repository: iic/GUI-Owl-7B
last_updated: 2025-08-27
license: mit
pipeline_tag: image-text-to-text
tasks:
  - image-text-to-text
model_type:
  - qwen2_5_vl
architectures:
  - Qwen2_5_VLForConditionalGeneration
base_model:
  - Qwen/Qwen2.5-VL-7B-Instruct
base_model_relation: finetune
parameters: 8.3B
tensor_type:
  - BF16
library_name:
  - safetensors
  - pytorch
frameworks:
  - Pytorch
language:
  - en
inference_backends:
  - "deploy_task vlm/text/emb"
  - "lmdeploy 0.9.1"
  - "sglang 0.5.2"
  - "vllm 0.9.2"
supports_inference: txt2txt
downloads: 9576
stars: 29
tags:
  - "arxiv:2508.15144"
---

# GUI-Owl-7B

> GUI-Owl-7B - iic 在 ModelScope 开源的模型。GUI-Owl is a model series developed as part of the Mobile-Agent-V3 project. It achieves state-of-the-art performance across a range of GUI automation benchmarks, including ScreenSpot-V2, ScreenSpot-Pro, OSWorld-G,…

iic/GUI-Owl-7B 是 ModelScope 魔搭社区上的 8.3B 参数image-text-to-text模型，采用 mit 许可，基于 Qwen/Qwen2.5-VL-7B-Instruct 构建，可用 deploy_task vlm/text/emb、lmdeploy 0.9.1、sglang 0.5.2 部署，并支持在线推理（txt2txt）。

- **Repository**: iic/GUI-Owl-7B
- **License**: mit
- **Tasks**: image-text-to-text
- **Parameters**: 8.3B
- **Base model**: Qwen/Qwen2.5-VL-7B-Instruct
- **Inference backends**: deploy_task vlm/text/emb, lmdeploy 0.9.1, sglang 0.5.2, vllm 0.9.2
- **Online inference**: txt2txt
- **Tags**: arxiv:2508.15144
- **Downloads**: 9576
- **Stars**: 29
- **Last updated**: 2025-08-27

Source: https://www.modelscope.cn/models/iic/GUI-Owl-7B

---

# GUI-Owl

<div align="center">
<img src=https://youke1.picui.cn/s1/2025/08/18/68a2f82fef3d4.png width="40%"/>
</div>

GUI-Owl is a model series developed as part of the Mobile-Agent-V3 project. It achieves state-of-the-art performance across a range of GUI automation benchmarks, including ScreenSpot-V2, ScreenSpot-Pro, OSWorld-G, MMBench-GUI, Android Control, Android World, and OSWorld. Furthermore, it can be instantiated as various specialized agents within the Mobile-Agent-V3 multi-agent framework to accomplish more complex tasks.

* **Paper**: [Paper Link](https://github.com/X-PLUG/MobileAgent/blob/main/Mobile-Agent-v3/assets/MobileAgentV3_Tech.pdf)
* **GitHub Repository**: https://github.com/X-PLUG/MobileAgent
* **Online Demo**: Comming soon

## Performance

### ScreenSpot-V2, ScreenSpot-Pro and OSWorld-G
<img src="https://github.com/X-PLUG/MobileAgent/blob/main/Mobile-Agent-v3/assets/screenspot_v2.jpg?raw=true" width="80%"/>
<img src="https://github.com/X-PLUG/MobileAgent/blob/main/Mobile-Agent-v3/assets/screenspot_pro.jpg?raw=true" width="80%"/>
<img src="https://github.com/X-PLUG/MobileAgent/blob/main/Mobile-Agent-v3/assets/osworld_g.jpg?raw=true" width="80%"/>

### MMBench-GUI L1, L2 and Android Control
<img src="https://github.com/X-PLUG/MobileAgent/blob/main/Mobile-Agent-v3/assets/mmbench_gui_l1.jpg?raw=true" width="80%"/>
<img src="https://github.com/X-PLUG/MobileAgent/blob/main/Mobile-Agent-v3/assets/mmbench_gui_l2.jpg?raw=true" width="80%"/>
<img src="https://github.com/X-PLUG/MobileAgent/blob/main/Mobile-Agent-v3/assets/android_control.jpg?raw=true" width="60%"/>

### Android World and OSWorld-Verified
<img src="https://github.com/X-PLUG/MobileAgent/blob/main/Mobile-Agent-v3/assets/online.jpg?raw=true" width="60%"/>

## Usage

Please refer to our cookbook.

## Deploy

We recommand deploy GUI-Owl-7B through vllm

This script has been validated on an A100 with 96 GB of VRAM.
```bash
PIXEL_ARGS='{"min_pixels":3136,"max_pixels":10035200}'
IMAGE_LIMIT_ARGS='image=2'
MP_SIZE=1
MM_KWARGS=(
    --mm-processor-kwargs $PIXEL_ARGS
    --limit-mm-per-prompt $IMAGE_LIMIT_ARGS
)

vllm serve $CKPT \
    --max-model-len 32768 ${MM_KWARGS[@]} \
    --tensor-parallel-size $MP_SIZE \
    --allowed-local-media-path '/' \
    --port 4243
```

If you want GUI-Owl to recieve more than two images, you could increase `IMAGE_LIMIT_ARGS` and reduce `max_pixels`.

For example:
```bash
PIXEL_ARGS='{"min_pixels":3136,"max_pixels":3211264}'
IMAGE_LIMIT_ARGS='image=5'
MP_SIZE=1
MM_KWARGS=(
    --mm-processor-kwargs $PIXEL_ARGS
    --limit-mm-per-prompt $IMAGE_LIMIT_ARGS
)

vllm serve $CKPT \
    --max-model-len 32768 ${MM_KWARGS[@]} \
    --tensor-parallel-size $MP_SIZE \
    --allowed-local-media-path '/' \
    --port 4243
```

## Citation
If you find our paper and model useful in your research, feel free to give us a cite.
```
@misc{ye2025mobileagentv3foundamentalagentsgui,
      title={Mobile-Agent-v3: Foundamental Agents for GUI Automation}, 
      author={Jiabo Ye and Xi Zhang and Haiyang Xu and Haowei Liu and Junyang Wang and Zhaoqing Zhu and Ziwei Zheng and Feiyu Gao and Junjie Cao and Zhengxi Lu and Jitong Liao and Qi Zheng and Fei Huang and Jingren Zhou and Ming Yan},
      year={2025},
      eprint={2508.15144},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2508.15144}, 
}
```
