---
title: eyes-4b
canonical_url: "https://www.modelscope.cn/models/gxcsoccer/eyes-4b"
md_url: "https://www.modelscope.cn/models/gxcsoccer/eyes-4b.md"
repository: gxcsoccer/eyes-4b
chinese_name: "得心 Eyes 4B"
last_updated: 2026-10-01
license: apache-2.0
pipeline_tag: image-text-to-text
tasks:
  - image-text-to-text
model_type:
  - qwen3_vl
architectures:
  - Qwen3VLForConditionalGeneration
base_model:
  - mPLUG/GUI-Owl-1.5-4B-Instruct
base_model_relation: quantized
parameters: 1.1B
tensor_type:
  - U32
  - BF16
library_name:
  - mlx
  - safetensors
  - pytorch
frameworks:
  - pytorch
language:
  - en
  - zh
downloads: 4
stars: 0
tags:
  - deskmind
  - mlx
  - gui-grounding
  - computer-use
  - screenspot
---

# eyes-4b

> eyes-4b - gxcsoccer 在 ModelScope 开源的模型。A 4B GUI grounding model: given a screenshot and a description of a UI element, it returns the point to click. DeskMind uses it on this Mac for apps that expose no accessibility tree.

gxcsoccer/eyes-4b 是 ModelScope 魔搭社区上的 1.1B 参数image-text-to-text模型，采用 apache-2.0 许可，基于 mPLUG/GUI-Owl-1.5-4B-Instruct 构建。

- **Repository**: gxcsoccer/eyes-4b
- **License**: apache-2.0
- **Tasks**: image-text-to-text
- **Parameters**: 1.1B
- **Base model**: mPLUG/GUI-Owl-1.5-4B-Instruct
- **Tags**: deskmind, mlx, gui-grounding, computer-use, screenspot
- **Downloads**: 4
- **Stars**: 0
- **Last updated**: 2026-10-01

Source: https://www.modelscope.cn/models/gxcsoccer/eyes-4b

---

# DeskMind Eyes 4B

A 4B GUI grounding model: given a screenshot and a description of a UI element, it returns the point to click.
DeskMind uses it on this Mac for apps that expose no accessibility tree.

一个 4B 的 GUI 定位模型：输入截图和对界面元素的描述，输出要点击的位置。
得心（DeskMind）在本机用它操作没有辅助功能结构的应用。

| Revision | Contents |
|---|---|
| `main` | bf16 weights (Transformers / vLLM), for reproducing the results below |
| `mlx-4bit` | MLX 4-bit build (group size 64) that the DeskMind app runs on Apple silicon |

## Results · 结果

The benchmarks were run with the bf16 weights, all on one platform: vLLM 0.19, the Qwen3-VL computer_use tool
prompt, native resolution and greedy decoding. Comparisons with other models are paired, item by item, on that platform.

用 bf16 权重评测，所有结果都在同一平台上得出：vLLM 0.19、Qwen3-VL computer_use tool prompt、原生分辨率、贪心解码。与其他模型的对比是在该平台上逐题配对进行的。

| Benchmark | GUI-Owl-1.5-4B (base) | **Eyes 4B** |
|---|---|---|
| ScreenSpot-Pro, no zoom, full 1581 | 64.8 | **67.7** |
| ScreenSpot-Pro, with zoom-in, full | 76.2 | **77.5** |
| ScreenSpot-v2, 1272 | 92.8 | **95.0** |

- On ScreenSpot-Pro (no zoom), Eyes 4B is +2.9 over the base (paired: +81 / −34, z = 4.4). It is +1.6 over KV-Ground-4B measured on the same platform (+68 / −42, z = 2.5).
- The `mlx-4bit` build as DeskMind deploys it (4-bit, screenshots scaled to ≤ 2 MP, `point_2d` prompt) has **no benchmark number**. Quantization and the smaller input may cost accuracy.

- ScreenSpot-Pro（不放大）上比基座高 2.9（配对 +81 / −34，z = 4.4）；比在同一平台实测的 KV-Ground-4B 高 1.6（+68 / −42，z = 2.5）。
- DeskMind 实际部署的 `mlx-4bit` 版本（4-bit、截图缩到 ≤ 2 MP、`point_2d` 提示词）**没有评测分数**。量化和更小的输入都可能降低准确率。

## Training · 训练

The base is GUI-Owl-1.5-4B-Instruct. On top of it, a LoRA (r = 32, on all language-model linear layers and
lm_head, vision tower frozen) was trained with GRPO and DAPO dynamic sampling, rewarded 1 for a click inside the
target box and 0 otherwise. The LoRA is merged into the weights published here.

1. DAPO on ShowUI-desktop and OS-Atlas desktop, with the OS-Atlas instructions rewritten into functional ones (see
   Disclosures).
2. Continued DAPO on a harder pool of GroundCUA records, 3.3K with a base pass rate in [1/8, 4/8], at 2.5 MP.

基座是 GUI-Owl-1.5-4B-Instruct。在其上训练了一个 LoRA（r = 32，作用于语言模型的全部线性层和 lm_head，视觉塔冻结），用 GRPO 和 DAPO 动态采样；点击落在目标框内奖励为 1，否则为 0。这里发布的权重已经合并了该 LoRA。

1. 在 ShowUI-desktop 和 OS-Atlas desktop 上做 DAPO。OS-Atlas 的指令改写成了功能描述式（见"说明"一节）。
2. 在更难的 GroundCUA 子集上继续 DAPO：3.3K 条基座通过率在 [1/8, 4/8] 之间的记录，训练分辨率 2.5 MP。

## Licence · 许可

The weights are released under **Apache-2.0**.

| Component | Licence |
|---|---|
| Base model GUI-Owl-1.5-4B-Instruct (mPLUG) | MIT: its notice is kept below |
| Qwen3-VL (the base of GUI-Owl-1.5) | Apache-2.0 |
| OS-Atlas data (OS-Copilot) | Apache-2.0 |
| GroundCUA (ServiceNow) | MIT |
| ShowUI-desktop (showlab) | **not stated**: see Disclosures |

权重以 **Apache-2.0** 发布。各组成部分的许可见上表：基座模型的 MIT 声明保留在本节末尾；ShowUI-desktop 的许可**未标明**，见"说明"一节。

> GUI-Owl-1.5 — Copyright (c) the mPLUG / X-PLUG authors. Released under the MIT License: permission is hereby granted,
> free of charge, to any person obtaining a copy of this software and associated documentation files, to deal in
> the Software without restriction, subject to the condition that the above copyright notice and this permission
> notice shall be included in all copies or substantial portions of the Software. THE SOFTWARE IS PROVIDED "AS IS",
> WITHOUT WARRANTY OF ANY KIND.

## Disclosures · 说明

- **ShowUI-desktop's licence is unconfirmed.** Its dataset card states none, and we have asked the authors
  ([showlab/ShowUI#104](https://github.com/showlab/ShowUI/issues/104)). If it turns out not to allow this use, a
  variant trained without it will replace these weights.
- **Rewritten instructions.** The OS-Atlas desktop instructions (accessibility names) were rewritten into functional
  instructions by a frontier model. That rewriting prompt used a handful of **ScreenSpot-Pro test instructions,
  text only**, as style examples: no screenshots, boxes or answers. The rewritten data was used for training, not for
  evaluation, but the prompt did see those test strings.
- **UI-Vision is not reported here.** GroundCUA shares its source with UI-Vision.

- **ShowUI-desktop 的许可尚未确认。** 数据卡上没有写许可，我们已经向作者询问（[showlab/ShowUI#104](https://github.com/showlab/ShowUI/issues/104)）。如果它不允许这种用途，会用不含它的数据重训一版，替换这里的权重。
- **改写过的指令。** OS-Atlas desktop 的指令（辅助功能名称）由一个前沿大模型改写成了功能描述式。改写用的提示词里，拿了几条 **ScreenSpot-Pro 测试集的指令（仅文字）**当风格示例，没有用截图、标注框或答案。改写后的数据只用于训练、不用于评测，但提示词确实见过这几条测试文字。
- **这里不报告 UI-Vision 分数。** 因为 GroundCUA 与 UI-Vision 同源。

## Use · 用法

Prompt (coordinates are 0–1000, relative to the image):

提示词（坐标为 0–1000，相对于图片）：

```
Locate the UI element in the screenshot that matches the description, and output its click point in JSON as
{"point_2d": [x, y]}, with coordinates in 0-1000 relative to the image.
Description: <instruction>
```

```bash
pip install mlx-vlm huggingface_hub
hf download deskmind/eyes-4b --revision mlx-4bit --local-dir eyes-4b-mlx
python -m mlx_vlm.generate --model eyes-4b-mlx --max-tokens 32 --image screenshot.png --prompt "<the prompt above>"
```

Integrity: `eyes-4b-bf16.sha256` on `main` and `eyes-4b-mlx4.sha256` on `mlx-4bit` list every file's SHA-256.

完整性校验：`main` 上的 `eyes-4b-bf16.sha256` 和 `mlx-4bit` 上的 `eyes-4b-mlx4.sha256` 列出了每个文件的 SHA-256。
