---
title: Eagle3-Qwen3-32B-zh
canonical_url: "https://www.modelscope.cn/models/Zjcxy-SmartAI/Eagle3-Qwen3-32B-zh"
md_url: "https://www.modelscope.cn/models/Zjcxy-SmartAI/Eagle3-Qwen3-32B-zh.md"
repository: Zjcxy-SmartAI/Eagle3-Qwen3-32B-zh
last_updated: 2025-12-16
license: "Apache License 2.0"
pipeline_tag: question-answering
tasks:
  - question-answering
  - text-generation
  - nli
model_type:
  - llama
architectures:
  - LlamaForCausalLMEagle3
library_name:
  - pytorch
frameworks:
  - Pytorch
downloads: 579
stars: 2
---

# Eagle3-Qwen3-32B-zh

> Eagle3-Qwen3-32B-zh - Zjcxy-SmartAI 在 ModelScope 开源的模型。Eagle3-Qwen3-32B-zh 基于开源的 Qwen3-32B 模型再训练所得，可用于 Eagle-3 投机解码算法，具备中/英文本加速能力，旨在加速语言大模型在解码阶段的推理速度。该模型使用中英文混合数据集进行训练，提高了中文接受率，适用于处理中/英文本的推理任务。

Zjcxy-SmartAI/Eagle3-Qwen3-32B-zh 是 ModelScope 魔搭社区上的question-answering、text-generation、nli模型，采用 Apache License 2.0 许可。

- **Repository**: Zjcxy-SmartAI/Eagle3-Qwen3-32B-zh
- **License**: Apache License 2.0
- **Tasks**: question-answering, text-generation, nli
- **Downloads**: 579
- **Stars**: 2
- **Last updated**: 2025-12-16

Source: https://www.modelscope.cn/models/Zjcxy-SmartAI/Eagle3-Qwen3-32B-zh

---

# Eagle3-Qwen3-32B-zh

## 💡 介绍

**Eagle3-Qwen3-32B-zh** 基于开源的 [**Qwen3-32B**](https://www.modelscope.cn/models/Qwen/Qwen3-32B) 模型再训练所得，可用于 Eagle-3 投机解码算法，具备**中/英文本加速能力**，旨在加速语言大模型在解码阶段的推理速度。该模型使用中英文混合数据集进行训练，提高了中文接受率，适用于处理中/英文本的推理任务。

为降低训练成本，我们基于 Eagle 和 SpecForge 框架，搭建了一套面向消费级显卡和一体机生产环境的 Eagle-3 训练工具。其可在 NVIDIA RTX 4090 一体机上运行端到端的训练流程，且优化了运行效率，约用**一周时间**完成此次训练。

值得说明的是，我们用约 **100k** 条数据训练所得的草稿模型，其加速性能优于当前已开源的同类权重（英文加速表现略优，中文加速显著提升）。经实测，在一体机 4 卡 NVIDIA RTX 4090 配置下，GSM8K 上 4 并发输出吞吐**最高可达 🚀 近 2 倍**的解码加速。

> 注：项目内配置文件基于 SGLang 框架进行配置，如若用于 vLLM 等其他框架，请自行修改相关配置。

## ⚙️ 训练配置

在硬件资源有限的条件下，我们基于 [**Eagle**](https://github.com/SafeAILab/EAGLE) 和 [**SpecForge**](https://github.com/sgl-project/SpecForge) 开源项目，搭建了一套面向消费级显卡的低成本 Eagle-3 权重训练工具，支持常用规模的大语言模型在一体机环境中的高效 Eagle-3 训练。

- 数据集：使用 ShareGPT-68K（英文）和 ShareGPT-Chinese-English-90k（中文）构建了 100K 条文本的训练数据集。训练数据集的 prompt 部分使用原数据集内容，output 部分则由 Qwen3-32B 模型重新生成。
- 训练环境：在 8 卡 24G 显存的 NVIDIA RTX 4090 显卡上展开训练，训练时长约一周。

## 🖥️ 推理命令

sglang 启动 eagle-3 算法服务

```shell
python3 -m sglang.launch_server \
--model-path Qwen/Qwen3-32B \
--speculative-algo EAGLE3 \
--speculative-draft Zjcxy-SmartAI/Eagle3-Qwen3-32B-zh \
--speculative-num-steps 5 \
--speculative-eagle-topk 4 \
--speculative-num-draft-tokens 16 \
--mem-fraction 0.7 \
--dtype float16 \
--served-model-name qwen \
--tp-size 4 \
--cuda-graph-max-bs 4
```

sglang 启动原始模型服务(用于对比实验)

```shell
 python -m sglang.launch_server \
--model-path Qwen/Qwen3-32B \
--mem-fraction 0.7 \
--dtype float16 \
--served-model-name qwen \
--tp-size 4 \
--cuda-graph-max-bs 4
```

## 🌟 性能测评
在 4 张 NVIDIA RTX 4090 上，我们基于 Eagle-3 算法，在以下数据集上对开源权重进行了系统性的测试：
- MT-bench ：该数据集包含了 80 条英文多轮高质量的对话，涵盖了不同领域。
- MT-bench-zh：中文版本的 MT-bench 数据集是通过[通义千问](https://www.tongyi.com/qianwen/)官网的服务处理后，再经过人工修正所得。
- C-Eval：从 C-Eval 数据集的每个类别中各抽取了 2 条数据，共构成了 104 条样本。
- GSM8K：从 GSM8K 数据集中随机抽取了 100 条数据。
> 注：所有测试均在 4 并发下进行，生成 token 数设置上限为 512。

为确保结论的充分性，我们做了以下实验实验：

### 🧪 实验一
设置每次推理进行 5 步投机解码迭代，保留 top-4 最高概率的 token 进行推测，最多选取 16 个草稿 token 进行并行验证。
```shell
--speculative-num-steps 5 \
--speculative-eagle-topk 4 \
--speculative-num-draft-tokens 16 \
```
<table>
  <thead>
    <tr>
        <th>&nbsp</th><th>&nbsp</th>
        <th colspan="2" style="text-align: center; vertical-align: middle;">MT-bench-zh</th>
        <th colspan="2" style="text-align: center; vertical-align: middle;">MT-bench</th>
        <th colspan="2" style="text-align: center; vertical-align: middle;">C-Eval</th>
        <th colspan="2" style="text-align: center; vertical-align: middle;">GSM8K</th>
    <tr><th>Temperature</th><th>Model</th><th>Speedup</th><th>τ</th><th>Speedup</th><th>τ</th><th>Speedup</th><th>τ</th><th>Speedup</th><th>τ</th></tr>
  </thead>
  <tbody>
    <!-- <tr><td colspan="12" style="text-align: center; vertical-align: middle;"><strong>Temperature=0</strong></td></tr> -->
    <tr><td rowspan="6"><strong>T=0</strong></td>
      <td>Zjcxy-SmartAI/Eagle3-Qwen3-32B</td><td><b>1.63x</b></td><td><b>3.37</b></td><td><b>1.73x</b></td><td><b>3.57</b></td><td><b>1.61x</b></td><td><b>3.3</b></td><td><b>1.91x</b></td><td><b>4.03</b></td></tr>
    <tr> 
      <tr><td>RedHatAI/Qwen3-32B-speculator.eagle3</td><td>0.69x</td><td>1.4</td><td>1.58x</td><td>3.24</td><td> 1.36x </td><td> 0.66 </td><td> 1.66x </td><td> 3.53 </td></tr>
  </tbody>
</table>


### 🧪 实验二
在实验一的基础上，基于 Qwen3-32B 官方推荐的采样参数配置，Temperature=0.6，TopP=0.95，TopK=20 和 MinP=0 。我们也进行了相关测试。

<table>
  <thead>
    <tr>
        <th>&nbsp</th><th>&nbsp</th>
        <th colspan="2" style="text-align: center; vertical-align: middle;">MT-bench-zh</th>
        <th colspan="2" style="text-align: center; vertical-align: middle;">MT-bench</th>
        <th colspan="2" style="text-align: center; vertical-align: middle;">C-Eval</th>
        <th colspan="2" style="text-align: center; vertical-align: middle;">GSM8K</th>
    <tr><th>Temperature</th><th>Model</th><th>Speedup</th><th>τ</th><th>Speedup</th><th>τ</th><th>Speedup</th><th>τ</th><th>Speedup</th><th>τ</th></tr>
  </thead>
  <tbody>
    <!-- <tr><td colspan="12" style="text-align: center; vertical-align: middle;"><strong>Temperature=0</strong></td></tr> -->
    <tr><td rowspan="6"><strong>T=0.6</strong></td>
      <td>Zjcxy-SmartAI/Eagle3-Qwen3-32B-zh</td><td><b>1.42x<b></td><td><b>3.27<b></td><td><b>1.5x</b></td><td><b>3.49</b></td><td><b>1.28x<b></td><td><b>3<b></td><td><b>1.67x</b></td><td><b>3.99</b></td></tr>
    <tr> </tr>
    <tr><td>RedHatAI/Qwen3-32B-speculator.eagle3</td><td> 0.62x</td><td> 1.39 </td><td>1.36x</td><td>3.15</td><td> - </td><td> - </td><td> 1.48x</td><td> 3.53</td></tr>
  </tbody>
</table>

## ✏️ 实验结论

通过对比，我们获得如下结论:
- 综合来看，在训练数据规模为 **100k** 的条件下，模型在各项评测中均展现出不错的性能。与当前同规模的开源权重相比，本模型的猜词能力和加速效果**在维持英文任务优势的基础上，在中文任务上取得了显著进步**。实验进一步表明：模型在中英文场景中的推理效率均可提升近 **1.5 倍**。


## 🔑 相关链接

Qwen3-32B 开源权重: https://www.modelscope.cn/models/Qwen/Qwen3-32B

Eagle 开源仓库：https://github.com/SafeAILab/EAGLE

SpecForce 训练框架：https://github.com/sgl-project/SpecForge
