---
title: Darwin-27B-ZTC-v2
canonical_url: "https://www.modelscope.cn/models/FINAL-Bench/Darwin-27B-ZTC-v2"
md_url: "https://www.modelscope.cn/models/FINAL-Bench/Darwin-27B-ZTC-v2.md"
repository: FINAL-Bench/Darwin-27B-ZTC-v2
last_updated: 2026-10-09
license: apache-2.0
pipeline_tag: text-classification
tasks:
  - text-classification
model_type:
  - qwen3_5_text
architectures:
  - Qwen3_5TextModel
parameters: 25.6B
tensor_type:
  - BF16
library_name:
  - safetensors
  - pytorch
frameworks:
  - pytorch
language:
  - en
downloads: 42
stars: 0
tags:
  - darwin
  - ztc
  - zero-token-classifier
  - decision-engine
  - system-one
  - s1mb
  - calibration
  - probabilistic-classification
---

# Darwin-27B-ZTC-v2

> Darwin-27B-ZTC-v2 - FINAL-Bench 在 ModelScope 开源的模型。A zero-token decision engine from the Darwin family, second version. Darwin-27B-ZTC-v2 reads a piece of state and a typed question (noul yes or no, choice one of N labels, score an ordered rubric) and…

FINAL-Bench/Darwin-27B-ZTC-v2 是 ModelScope 魔搭社区上的 25.6B 参数text-classification模型，采用 apache-2.0 许可。

- **Repository**: FINAL-Bench/Darwin-27B-ZTC-v2
- **License**: apache-2.0
- **Tasks**: text-classification
- **Parameters**: 25.6B
- **Tags**: darwin, ztc, zero-token-classifier, decision-engine, system-one, s1mb, calibration, probabilistic-classification
- **Downloads**: 42
- **Stars**: 0
- **Last updated**: 2026-10-09

Source: https://www.modelscope.cn/models/FINAL-Bench/Darwin-27B-ZTC-v2

---

# Darwin-27B-ZTC-v2

<a href="https://huggingface.co/spaces/hotchpotch/S1MB-leaderboard"><img src="https://img.shields.io/badge/S1MB_leaderboard-%231_(Borda_89.58)-7c3aed?style=for-the-badge"></a>

**A zero-token decision engine from the Darwin family, second version.** Darwin-27B-ZTC-v2 reads a piece of state and a typed question (`noul` yes or no, `choice` one of N labels, `score` an ordered rubric) and returns a probability for every option. It uses one forward pass per question and generates no tokens.

v2 adds control and workflow decisions (games, drones, retrieval control, entity alignment, customer and incident workflows) on top of [FINAL-Bench/Darwin-27B-ZTC](https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC).

## Results

### S1MB (System One Mosaic Benchmark), english-v1: #1

**#1 on the [S1MB leaderboard](https://huggingface.co/spaces/hotchpotch/S1MB-leaderboard)** under both of its rankings: the default **Borda score** and **Task Avg**. All 137 benchmarks complete, measured with the S1MB evaluator and its `autojev` adapter at revision `a8f283a`, then validated and merged by the S1MB maintainer ([results](https://huggingface.co/datasets/hotchpotch/s1mb-result/tree/main/FINAL-Bench__Darwin-27B-ZTC-v2)).

The numbers below follow the leaderboard's own code (`viewer/src/lib/borda.ts`, `types.ts`) applied to the published results file of 2026-10-09 02:44 UTC, over the 102 models with complete results.

- **Borda score** (default sort): rank the models on every benchmark, give 100 points to first place and 0 to last, and average over the 137 benchmarks.
- **Task Avg**: mean baseline-adjusted score (0 to 100) within each task, then the mean of Noul, Choice and Score.

| Rank | Model | Borda score | Task Avg | Noul (59) | Choice (57) | Score (21) |
|---|---|---|---|---|---|---|
| **1** | **Darwin-27B-ZTC-v2** | **89.58** | **66.46** | **67.66** | **71.51** | **60.21** |
| 2 | openjev/openjev | 87.50 | 62.60 | 65.22 | 67.74 | 54.83 |
| 3 | denis-pplx/AutoJev-27B | 87.07 | 60.80 | 64.96 | 68.11 | 49.34 |
| 4 | caiovicentino1/Eikos-27B | 85.43 | 59.86 | 64.14 | 67.61 | 47.83 |
| 5 | TypeSafe Jev 1.13 | 85.05 | 59.59 | 64.63 | 67.22 | 46.92 |

`qasper-noul-test-v1` was run with `--context-limit 32768` (12 of its decisions exceed the default 8192 tokens; nothing is truncated), recorded in the result metadata.

**Training overlap (disclosed).** v2 was trained on the `train` split of [ZefanCai/Open-Jev](https://huggingface.co/datasets/ZefanCai/Open-Jev) (CC0). S1MB includes 22 benchmarks built from the `test` split of the same Open-Jev tasks. We used no S1MB test case: an exact-match filter over every S1MB test state, query and document removed zero training rows. v2 was not trained on any S1MB data or on Typed Decisions data.

## How it was made

1. Start from Darwin-27B-ZTC (v1).
2. Continue full-weight training on the Open-Jev `train` split mixed with v1's original training data, selecting the checkpoint on held-out development rows only.
3. Average the weights of v1 and the continued model (50/50). The average keeps v1's skills on its original tasks and adds the new control and workflow skills.

## Files and usage

The layout is the same as v1: backbone weights (BF16, about 54 GB), `readout.safetensors`, `decision_config.json`, tokenizer, and the inference code in `autojev/` (from [autojev](https://github.com/denis-pplx/autojev), MIT, with text-only backbone support added).

- `ztc_server.py`: a `POST /v1/systemone` server (`ZTC_MODEL=<path> PORT=8000 python ztc_server.py`).
- `ztc_engine.py`: an in-process engine for the Decision Index kit.

```python
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("FINAL-Bench/Darwin-27B-ZTC-v2")
sys.path.insert(0, path)
from autojev.model import DecisionModel

model = DecisionModel(checkpoint=path)
row = {"state": {"ticket": "Charged twice for one order."},
       "question": {"type": "choice", "instructions": "What should support do?",
                    "criteria": {"refund": "Refund the duplicate charge.", "escalate": "Send to billing.", "close": "No action."}}}
print(model.predict([row])[0])
```

## Citation

```bibtex
@misc{darwin27bztcv2,
  title  = {Darwin-27B-ZTC-v2: a zero-token decision engine},
  author = {VIDRAFT and FINAL-Bench},
  year   = {2026},
  url    = {https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC-v2}
}
```
