---
title: Qwen3.6-27B-DSV4Pro-Thinking-Distill-NVFP4
canonical_url: "https://www.modelscope.cn/models/Merkyor/Qwen3.6-27B-DSV4Pro-Thinking-Distill-NVFP4"
md_url: "https://www.modelscope.cn/models/Merkyor/Qwen3.6-27B-DSV4Pro-Thinking-Distill-NVFP4.md"
repository: Merkyor/Qwen3.6-27B-DSV4Pro-Thinking-Distill-NVFP4
last_updated: 2026-06-28
model_type:
  - qwen3_5
architectures:
  - Qwen3_5ForConditionalGeneration
parameters: 20.3B
tensor_type:
  - F32
  - BF16
  - F8_E4M3
  - U8
library_name:
  - safetensors
downloads: 547
stars: 2
---

# Qwen3.6-27B-DSV4Pro-Thinking-Distill-NVFP4

> Qwen3.6-27B-DSV4Pro-Thinking-Distill-NVFP4 - Merkyor 在 ModelScope 开源的模型。Qwen3.6-27B DSV4Pro 思维蒸馏 — NVFP4（多档 + 官方 MTP）

- **Repository**: Merkyor/Qwen3.6-27B-DSV4Pro-Thinking-Distill-NVFP4
- **Parameters**: 20.3B
- **Downloads**: 547
- **Stars**: 2
- **Last updated**: 2026-06-28

Source: https://www.modelscope.cn/models/Merkyor/Qwen3.6-27B-DSV4Pro-Thinking-Distill-NVFP4

---

# Qwen3.6-27B DSV4Pro 思维蒸馏 — NVFP4（多档 + 官方 MTP）

本仓为 [27B DSV4Pro 思维蒸馏](https://modelscope.cn/models/Merkyor/Qwen3.6-27B-DSV4Pro-Thinking-Distill) 的 **NVFP4 量化发布**，含**多档量化 + 官方 nextn MTP 头**。NVFP4 用 **W4A16 / W4A4（/ 未来 W4A8）+ MTP 一并兼顾质量、速度、并发**。

## 量化档位

| 档位 | 目录 | GPQA-D 198 | MMLU-500 | 大小 | 定位 |
|------|------|-----------|----------|------|------|
| **W4A16**（默认，根目录） | `/` | **82.83%** | 87.80% | 29 GB | 质量优先 |
| **W4A4** | `w4a4/` | 77.27% | **91.40%** | 20 GB | 速度 / 显存优先 |
| W4A8 | *待引擎支持* | ≈ W4A16（预期） | — | ~19 GB | vLLM 暂不支持，见末尾 |

- **W4A16**：MLP-focused —— 仅量化 language MLP 权重（`gate/up/down`），attention、Mamba-GDN、视觉、embeddings、`lm_head`、norm 保 BF16。质量最高。
- **W4A4**：量化 MLP + attention + GDN 投影，仅 conv1d 保 BF16。更小更快，质量略降但 MMLU 反超（thinking-on 口径）。
- 两档均**内置官方 MTP 头**（`mtp.safetensors`），vLLM 开启即加速。

## MTP 投机加速（官方 nextn 头）

> 所有速度 **@ RTX PRO 6000 Blackwell (sm120) + vLLM 0.23**（区别于 Spark/GB10 的 GGUF 速度，勿混淆）。

| 档位 | 推荐 `num_speculative_tokens` | 单流 tok/s（MTP / 非MTP） | 加速 | accept |
|------|------------------------------|--------------------------|------|--------|
| W4A16 | **4** | 93.7 / 45 | 2.08× | 0.83 |
| W4A4 | **2** | 111 / 63 | 1.76× | 0.87 |

**MTP 并发同样正收益**（W4A4 `N=2` @R6000，聚合 tok/s）：

| 并发 | 非MTP | MTP | 提升 |
|------|-------|-----|------|
| 1 | 65 | 112 | +72% |
| 4 | 235 | 371 | +58% |
| 8 | 453 | 672 | +48% |
| 16 | 825 | 1192 | +44% |

MTP 优势随并发递减、但 **c1–c16 全程为正**（NVFP4 权重小 memory-bound 区间宽 + sm120 算力强，临界并发 > 16）。仅在 GPU 完全饱和的极高并发下才可能转负——按实际负载自测即可。

## 加载

### W4A16（默认 / 根目录）
```bash
# 纯质量
vllm serve Merkyor/Qwen3.6-27B-DSV4Pro-Thinking-Distill-NVFP4 \
  --quantization modelopt --max-model-len 32768

# 开 MTP（单流 / 低并发，推荐 N=4）
vllm serve Merkyor/Qwen3.6-27B-DSV4Pro-Thinking-Distill-NVFP4 \
  --quantization modelopt --max-model-len 32768 \
  --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":4}'
```

### W4A4（`w4a4/` 子目录）
```bash
# 先下载整仓，再指向 w4a4 子目录
modelscope download --model Merkyor/Qwen3.6-27B-DSV4Pro-Thinking-Distill-NVFP4 --local_dir ./nvfp4
vllm serve ./nvfp4/w4a4 \
  --quantization modelopt --max-model-len 32768 \
  --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":2}'
```

### 推理参数（思考模式，默认）
`temperature=0.6`、`top_p=0.95`、`max_tokens=32000`。

### SGLang
serve 用 `--quantization modelopt_fp4`；MTP 用 `--speculative-algorithm NEXTN --speculative-num-steps N`（sm121/GB10 需 `SPEC_V2=1` + mamba/flashinfer flag）。**SGLang 的 num-steps 最优值与 vLLM 可能不同，待 Spark 实测补充。**

## 硬件
- NVIDIA FP4 Tensor Core：**Blackwell SM120+**（RTX PRO 6000、B200）
- 显存：W4A16 ~32 GB+ / W4A4 ~24 GB+
- vLLM ≥ 0.8（quant_algo 支持 `NVFP4` / `W4A16_NVFP4`）或支持 `modelopt_fp4` 的 SGLang

## 评测口径
temp 0.6 / top_p 0.95 / thinking-on / max_tokens 32000 / ctx 36864；GPQA-Diamond 198 全集 + MMLU-500 5-shot。质量分与硬件无关；速度数标注 @R6000。

## W4A8（待引擎支持）
W4A8（weight NVFP4 + act FP8）已可用 ModelOpt `W4A8_NVFP4_FP8_CFG` 量化，**质量预期 ≈ W4A16**（act FP8 近无损），R6000 硬件原生支持。但 **vLLM 0.23 的 ModelOpt loader 暂不接受 `W4A8_NVFP4_FP8` quant_algo**（仅 FP8/NVFP4/W4A16_NVFP4/MXFP8/MIXED_PRECISION），故暂不发布。待 vLLM 支持后将补 `w4a8/` 子目录。

## 相关仓库
| 仓库 | 用途 |
|------|------|
| [基础模型 BF16](https://modelscope.cn/models/Merkyor/Qwen3.6-27B-DSV4Pro-Thinking-Distill) | 完整精度 |
| [GGUF + MTP](https://modelscope.cn/models/Merkyor/Qwen3.6-27B-DSV4Pro-Thinking-Distill-GGUF) | llama.cpp 部署 |
| **NVFP4 多档 + MTP（本仓）** | vLLM / SGLang，质量+速度+并发 |

---

# Qwen3.6-27B DSV4Pro Thinking Distill — NVFP4 (multi-tier + official MTP)

NVFP4 quantization with **multiple tiers + official nextn MTP head**, balancing quality, speed, and concurrency via **W4A16 / W4A4 (/ future W4A8) + MTP**.

## Tiers

| Tier | Path | GPQA-D 198 | MMLU-500 | Size | Focus |
|------|------|-----------|----------|------|-------|
| **W4A16** (default, root) | `/` | **82.83%** | 87.80% | 29 GB | quality |
| **W4A4** | `w4a4/` | 77.27% | **91.40%** | 20 GB | speed / VRAM |
| W4A8 | *pending engine support* | ≈ W4A16 (expected) | — | ~19 GB | vLLM TBD |

- **W4A16**: MLP-focused (only MLP weights to NVFP4; attention/Mamba/vision/embeddings/`lm_head`/norm stay BF16).
- **W4A4**: MLP + attention + GDN quantized, conv1d kept BF16.
- Both ship the official **MTP head**.

## MTP (official nextn head) — all speeds @ RTX PRO 6000 (sm120), vLLM 0.23

| Tier | Recommended `num_speculative_tokens` | tok/s (MTP / no-MTP) | speedup | accept |
|------|--------------------------------------|----------------------|---------|--------|
| W4A16 | **4** | 93.7 / 45 | 2.08× | 0.83 |
| W4A4 | **2** | 111 / 63 | 1.76× | 0.87 |

**MTP stays net-positive under concurrency** (W4A4 N=2, aggregate tok/s): c1 +72% / c4 +58% / c8 +48% / c16 +44%. Gain decreases with concurrency but stays positive through c16 (NVFP4 is memory-bound-wide + sm120 is compute-strong; crossover concurrency > 16). Only turns negative at extreme saturation — benchmark for your load.

## Load

```bash
# W4A16 (root) + MTP
vllm serve Merkyor/Qwen3.6-27B-DSV4Pro-Thinking-Distill-NVFP4 --quantization modelopt \
  --max-model-len 32768 --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":4}'

# W4A4 (w4a4/ subdir) + MTP
modelscope download --model Merkyor/Qwen3.6-27B-DSV4Pro-Thinking-Distill-NVFP4 --local_dir ./nvfp4
vllm serve ./nvfp4/w4a4 --quantization modelopt \
  --max-model-len 32768 --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":2}'
```

Inference: `temperature=0.6`, `top_p=0.95`, `max_tokens=32000` (thinking mode). Hardware: Blackwell SM120+ (RTX PRO 6000 / B200).

## Claude Code（实验性 / experimental）

Claude Code 支持目前属于实验性接入。Claude Code 需要兼容 Anthropic `/v1/messages` 的服务端和稳定的工具调用；直接使用 OpenAI-compatible chat endpoint 可能无法正常工作，需要桥接或兼容运行时。

本模型使用 Qwen3 XML 工具调用格式（`<tool_call><function=name><parameter=...>`）。vLLM 路线：

```bash
vllm serve /path/to/model \
  --served-model-name qwen36-27b-distill \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml
```

Claude Code 指向 `served-model-name`（不要用带 `/` 的 HF repo id）；或使用 LM Studio 0.4.1+（内置 Claude Code `/v1/messages`）。参考 [vLLM Claude Code](https://docs.vllm.ai/en/stable/serving/integrations/claude_code/) · [LM Studio](https://lmstudio.ai/blog/claudecode)。

---

## Claude Code (experimental)

Claude Code support is experimental. It needs an Anthropic-compatible `/v1/messages` endpoint and stable tool calling; a direct OpenAI-compatible chat endpoint may not work and requires a bridge or compatible runtime. The model uses the Qwen3 XML tool format — on vLLM use `--tool-call-parser qwen3_xml` (command above). Point Claude Code at the `served-model-name` (no slashes), or use LM Studio 0.4.1+ (built-in Claude Code `/v1/messages`).
