---
title: brain-4b
canonical_url: "https://www.modelscope.cn/models/gxcsoccer/brain-4b"
md_url: "https://www.modelscope.cn/models/gxcsoccer/brain-4b.md"
repository: gxcsoccer/brain-4b
chinese_name: "得心 Brain 4B"
last_updated: 2026-10-01
license: apache-2.0
pipeline_tag: text-classification
tasks:
  - text-classification
model_type:
  - qwen3_5
architectures:
  - Qwen3_5ForConditionalGeneration
base_model:
  - Qwen/Qwen3.5-4B
base_model_relation: quantized
parameters: 1.2B
tensor_type:
  - U32
  - BF16
library_name:
  - mlx
  - safetensors
language:
  - en
  - zh
downloads: 6
stars: 0
tags:
  - computer-use
  - agents
  - decision-model
  - calibration
  - mlx
  - apple-silicon
  - local-llm
  - deskmind
---

# brain-4b

> brain-4b - gxcsoccer 在 ModelScope 开源的模型。DeskMind Brain 4b · 得心

gxcsoccer/brain-4b 是 ModelScope 魔搭社区上的 1.2B 参数text-classification模型，采用 apache-2.0 许可，基于 Qwen/Qwen3.5-4B 构建。

- **Repository**: gxcsoccer/brain-4b
- **License**: apache-2.0
- **Tasks**: text-classification
- **Parameters**: 1.2B
- **Base model**: Qwen/Qwen3.5-4B
- **Tags**: computer-use, agents, decision-model, calibration, mlx, apple-silicon, local-llm, deskmind
- **Downloads**: 6
- **Stars**: 0
- **Last updated**: 2026-10-01

Source: https://www.modelscope.cn/models/gxcsoccer/brain-4b

---

# DeskMind Brain 4b · 得心

**得心，应手。** Brain is the decision model of [DeskMind](https://github.com/deskmind-ai), open-source models and
tools that let an agent see your screen, decide the next step and act on your own Mac. It answers System One–style
typed questions (which operation next, which element, is the goal met, should I ask the user first) and returns
calibrated probabilities read from answer-letter logits: nothing is generated or parsed. The server is wire-compatible
with `POST /v1/systemone`.

This model is the **strong tier** of the Brain router: it answers the steps that [brain-0.8b](https://huggingface.co/deskmind/brain-0.8b) escalates (about 70% of steps in this release). It can also be served on its own.

- **Release:** G18b, revision `g18b-q8` (8-bit MLX, prompt format 3). `main` follows the current release.
- **Download:** 4.5 GB.
- **Code and docs:** [deskmind-ai/brain](https://github.com/deskmind-ai/brain). Keep `deskmind.json` next to the
  weights: it records the prompt format the model was trained with.

## Revisions

| Revision | What it is |
|---|---|
| `g18b-q8` | **Current release.** G18b, router threshold 0.96; used by the DeskMind app v0.3.0. |
| `g17-q8` | Test build, not a release. |
| `g14-q8` | Earlier release (G14, prompt format 3, router threshold 0.94); used by the DeskMind app v0.2.0. |
| `g13-q8` | Earlier release (G13, prompt format 3). |
| `v7b-q8` | First published build (checkpoint G11b, prompt format 2). |

## Use

```bash
git clone https://github.com/deskmind-ai/brain && cd brain
uv sync --extra mlx
uv run hf download deskmind/brain-0.8b --revision g18b-q8 --local-dir models/brain-0.8b
uv run hf download deskmind/brain-4b --revision g18b-q8 --local-dir models/brain-4b
uv run deskmind-brain-serve --predictor mlx:models/brain-0.8b --escalate-to mlx:models/brain-4b \
  --two-stage --port 8796
```

The 4B can also answer every step on its own:

```bash
uv run deskmind-brain-serve --predictor mlx:models/brain-4b --port 8793 --two-stage
```

The 4B is 4.5 GB to download and the 0.8B 0.8 GB. If `hf download` fails with `CAS Client Error`, retry with
`HF_HUB_DISABLE_XET=1` in front of the command. In mainland China, ModelScope carries the same files:

```bash
uvx modelscope download --model gxcsoccer/brain-4b --revision g18b-q8 --local-dir models/brain-4b
```

Each reply carries a `routing` record: who answered, and why. Full walkthrough:
[deskmind-ai/brain](https://github.com/deskmind-ai/brain#quick-start).

## Results

All numbers are our own runs; method and full tables are in
[docs/results.md](https://github.com/deskmind-ai/brain/blob/main/docs/results.md).

**Real macOS desktop**, [bench](https://github.com/deskmind-ai/bench) suite v25, 13 sandbox tasks × 3 runs, strict
pass, run through the DeskMind app on an M4 Pro (48 GB):

| config | pass | false "done" | decision time p50 / p95 |
|---|---|---|---|
| **Router G18b** (0.8B → 4B, 8-bit, threshold 0.96) | **39/39** | **0** | 2.85 / 9.82 s (208 decisions) |
| Router G14 (earlier release, threshold 0.94) | 36/39 | 0 | 0.57 / 5.25 s |

- Steps the 0.8B answers itself take p50 0.48 s / p95 0.66 s; steps escalated to the 4B take p50 3.6 s / p95 9.8 s.
  About 70% of steps escalate.
- The G18b runs had the app's optional checks and notes off; there were no environment errors and no no-progress
  loops.
- 39 runs over 13 tasks is a small sample, and runs cluster by task (a task tends to pass 3/3 or 0/3).
- On the earlier suite v23, the G14 router passed 35/38 (92%) and Jev (TypeSafe, hosted) 33/38 (87%). There is no
  Jev run on v25.

**JevBench v1.4.2, public set** (231 items, the board's `public_accuracy` column), run locally with the official
`jevbench.cli`:

| | easy (48) | original (72) | hard (111) | public (231) |
|---|---|---|---|---|
| **Brain 4B, G18b** | 48 | 69 | 76 | **0.835** |
| Brain 0.8B, G18b | 48 | 56 | 63 | 0.723 |
| Router G18b (threshold 0.96) | 48 | 65 | 71 | 0.797 |
| Brain 4B, G14 | 48 | 67 | 85 | 0.866 |

- Public items only. The board's headline JevBench Score also weighs 308 sealed items, calibration, speed and cost;
  we have not been scored on the sealed set.
- Calibration of the G18b 4B: Brier 0.269, ECE 0.089 (G14: 0.230, 0.079).
- Contamination check: the G18b training mix (100,703 items) shares no word 13-gram and no option set with the 231
  public items (0/231).

## Limitations

- **Speed:** the G18b 0.8B's confidences sit in a narrow band (about 0.94–0.97), so at threshold 0.96 most steps go
  to the 4B and a typical decision takes about 3 s, slower than hosted models.
- **General judgement:** against G14, the 4B dropped on JevBench's hard tier (85 → 76 of 111), mostly
  temporal/numeric, hard-judgement and multi-hop items. G18b was kept for its real-desktop reliability.
- **Saying "done":** on the Chinese exact-text task the file was right in all 3 runs, but the model never said "done"
  and used the full 20-step budget. The grader checks the final state, so these count as passes.
- **Scope:** trained and tested on macOS Finder and TextEdit sandbox tasks plus web and form decisions; untested
  elsewhere.
- **Probabilities are not guarantees:** a probability is the model's own weighting of the options, not proof that the
  step is right.

## Training

LoRA distillation on Qwen/Qwen3.5-4B (KL to teacher distributions plus cross-entropy to labels), merged and quantized to 8
bits. Desktop data comes from DAgger in sandboxed macOS tasks, with every visited state labelled by a desktop oracle,
plus counterexamples that break label shortcuts. Web, form and evidence items come from public datasets and synthetic
tasks, generated and labelled with hosted frontier-model teachers. No Jev outputs were used as labels, and the data
contains no real user data. Details:
[docs/training.md](https://github.com/deskmind-ai/brain/blob/main/docs/training.md).

---

**中文：** 得心（DeskMind）Brain 的强档：回答 0.8B 交上来的步骤（本版约占 70%），也可以单独运行。给电脑操作 agent 的每一步做带类型、带把握程度的决策，用 MLX 在 Apple Silicon 本地运行。
- **当前发布版：** G18b，版本 `g18b-q8`（8 位 MLX）；DeskMind app v0.3.0 使用此版本。
- **真机成绩：** bench v25，13 个沙箱任务各跑 3 轮，通过 DeskMind app 运行，39/39 通过，没做完就说完成 0 次。0.8B
  直接回答的步骤中位 0.48 秒；约 70% 的步骤交给 4B，中位 3.6 秒。
- **JevBench v1.4.2 公开题（231 道）：** 4B 0.835，0.8B 0.723，路由 0.797；训练数据与公开题无重合（0/231）。
- **下载：** `hf download` 加 `--revision g18b-q8`；国内可用 ModelScope：
  `uvx modelscope download --model gxcsoccer/brain-4b --revision g18b-q8 --local-dir models/brain-4b`。
- 详见 [deskmind-ai/brain](https://github.com/deskmind-ai/brain/blob/main/README.zh-CN.md)。

## License

Apache-2.0 (see `LICENSE` and `NOTICE`). Fine-tuned from Qwen/Qwen3.5-4B (Copyright Alibaba Cloud, Apache-2.0). The DeskMind
name, 得心 and the logo are not covered by this licence.
