---
title: Qwen3.6-35B-A3B-MTP-Donor
canonical_url: "https://www.modelscope.cn/models/HereIsMark/Qwen3.6-35B-A3B-MTP-Donor"
md_url: "https://www.modelscope.cn/models/HereIsMark/Qwen3.6-35B-A3B-MTP-Donor.md"
repository: HereIsMark/Qwen3.6-35B-A3B-MTP-Donor
chinese_name: Qwen3.6-35B-A3B-MTP-Donor
last_updated: 2026-05-18
library_name:
  - gguf
downloads: 89
stars: 0
---

# Qwen3.6-35B-A3B-MTP-Donor

> Qwen3.6-35B-A3B-MTP-Donor - HereIsMark 在 ModelScope 开源的模型。MTP Donor for Qwen3.6-35B-A3B

- **Repository**: HereIsMark/Qwen3.6-35B-A3B-MTP-Donor
- **Downloads**: 89
- **Stars**: 0
- **Last updated**: 2026-05-18

Source: https://www.modelscope.cn/models/HereIsMark/Qwen3.6-35B-A3B-MTP-Donor

---

# MTP Donor for Qwen3.6-35B-A3B

Multi-Token Prediction (MTP) donor file for **Qwen3.6-35B-A3B**.

**Edited: 2026/5/17 Since llama.cpp has already merged the MTP branch, this model is no longer relevant (although it still works).Please DO NOT USE this model ANYMORE**

> **This donor is ONLY compatible with Qwen3.6-35B-A3B.**  
> Do NOT use with other model sizes (122B, 27B, etc.) or architectures.

## What is this?

A lightweight GGUF file containing only the MTP (Multi-Token Prediction) layer tensors extracted from a full Qwen3.6-35B-A3B model. Instead of downloading the full model (tens of GB) just to get MTP support, you can download this donor and inject it into your existing GGUF.

## About the script

The `convert.py` script in this repo works with **any GGUF that contains MTP layers**. You can use it to extract and share MTP donors for other models too:

```bash
# Extract MTP from any model
python convert.py extract any-model-with-mtp.gguf donor.gguf

# Merge donor into base model
python convert.py merge base.gguf donor.gguf output.gguf
```

It has been tested with Qwen3.5-35B, Qwen3.6-27B, and Qwen3.5-122B-A10B — all work the same way. The donor preserves the exact quantization type and per-row metadata of the original tensors.

## Usage

### Prerequisites

You need a **custom build** of llama.cpp with MTP support:

- **Branch/PR**: https://github.com/ggml-org/llama.cpp/pull/22673
- You **must compile from source** — pre-built binaries do not include MTP support.

### Merge Script

Merge the MTP donor into your existing GGUF:
```bash
python convert.py merge your-model.gguf Qwen3.6-35B-A3B-MTP-Donor-Q8_0.gguf output.gguf
```

3. Run inference with the custom llama.cpp build, enabling MTP:
```bash
./llama-cli -m output.gguf -p "Your prompt" -n 256 --spec-type mtp --spec-draft-n-max 2
```

For the server:
```bash
./llama-server -m output.gguf --spec-type mtp --spec-draft-n-max 2
```

## Files

| File | Description |
|------|-------------|
| `Qwen3.6-35B-A3B-MTP-Donor-Q8_0.gguf` | MTP tensors Q8_0 (~897.5 MB, recommended) |

Q8_0 is faster than BF16 due to lower memory bandwidth, and the quantization error is corrected by the trunk model's rejection sampling.

## Credits

- [llama.cpp MTP PR #22673](https://github.com/ggml-org/llama.cpp/pull/22673)
- [AzerbaijanNyan](https://www.reddit.com/r/LocalLLaMA/comments/1t6r1ny/extracted_mtp_tensor_ggufs_smaller_donor_models/) — original extraction method
