---
title: Qwen3-VL-8B-Instruct-Eagle3
canonical_url: "https://www.modelscope.cn/models/MNN/Qwen3-VL-8B-Instruct-Eagle3"
md_url: "https://www.modelscope.cn/models/MNN/Qwen3-VL-8B-Instruct-Eagle3.md"
repository: MNN/Qwen3-VL-8B-Instruct-Eagle3
last_updated: 2025-11-14
license: apache-2.0
pipeline_tag: text-generation
tasks:
  - text-generation
model_type:
  - llama
architectures:
  - LlamaForCausalLMEagle3
base_model:
  - Qwen/Qwen3-VL-8B-Instruct
base_model_relation: finetune
parameters: 399.7M
tensor_type:
  - BF16
  - BOOL
  - I64
library_name:
  - safetensors
language:
  - zh
  - en
inference_backends:
  - "deploy_task text"
  - "sglang 0.5.2"
downloads: 6970
stars: 1
tags:
  - speculative-decoding
  - eagle
  - qwen-vl
---

# Qwen3-VL-8B-Instruct-Eagle3

> Qwen3-VL-8B-Instruct-Eagle3 - MNN 在 ModelScope 开源的模型。EAGLE-3 Draft Model for Qwen3-VL-8B-Instruct

MNN/Qwen3-VL-8B-Instruct-Eagle3 是 ModelScope 魔搭社区上的 399.7M 参数text-generation模型，采用 apache-2.0 许可，基于 Qwen/Qwen3-VL-8B-Instruct 构建，可用 deploy_task text、sglang 0.5.2 部署。

- **Repository**: MNN/Qwen3-VL-8B-Instruct-Eagle3
- **License**: apache-2.0
- **Tasks**: text-generation
- **Parameters**: 399.7M
- **Base model**: Qwen/Qwen3-VL-8B-Instruct
- **Inference backends**: deploy_task text, sglang 0.5.2
- **Tags**: speculative-decoding, eagle, qwen-vl
- **Downloads**: 6970
- **Stars**: 1
- **Last updated**: 2025-11-14

Source: https://www.modelscope.cn/models/MNN/Qwen3-VL-8B-Instruct-Eagle3

---

<!-- 语言切换 / Language Toggle -->
<div align="center">
<a href="#-en">English</a> | <a href="#-zh-cn">中文</a>
</div>

<!-- 英文版 README -->
<div id="-en">

# EAGLE-3 Draft Model for Qwen3-VL-8B-Instruct

## Model Overview

This repository contains an **EAGLE-3 style draft model** specifically trained to accelerate the inference of the `Qwen3-VL-8B-Instruct` large language model.

This is **not a standalone model**. It must be used in conjunction with its corresponding base model (`Qwen3-VL-8B-Instruct`) within a speculative decoding framework to achieve significant speedups in text generation.

- **Base Model:** `Qwen3-VL-8B-Instruct`
- **Model Architecture:** EAGLE-3 (Speculative Decoding Draft Model)
- **Primary Benefit:** Accelerates text generation throughput by 1.5x to 2.5x without compromising the generation quality of the base model.

## What is EAGLE?

EAGLE (Extrapolative A* Generative Language Engine) is an advanced speculative decoding method. It uses a small draft model to generate a sequence of draft tokens in parallel. These tokens are then verified by the larger, more powerful base model in a single forward pass. If the draft is accepted, the generation process advances multiple steps at once, leading to a substantial increase in speed.

This model serves as the "draft model" in this process. Its average acceptance length (`acc_length`) on standard benchmarks is approximately **1.87 tokens** (with 4 draft tokens), meaning on average, it helps the base model advance nearly 2 tokens per verification step.

## Performance

This model was evaluated on a diverse set of benchmarks. The `acc_length` (average number of accepted draft tokens) indicates the efficiency of the acceleration. A higher value is better.

| Benchmark  | `acc_length` (num_draft_tokens=4) | `acc_length` (num_draft_tokens=8) |
| :--------- | :-------------------------------: | :-------------------------------: |
| math500    |               2.30                |               2.55                |
| humaneval  |               2.23                |               2.43                |
| gsm8k      |               2.05                |               2.15                |
| ceval      |               1.84                |               1.93                |
| cmmlu      |               1.81                |               1.88                |
| mtbench    |               1.76                |               1.83                |
| **Average**|             **~2.00**             |             **~2.13**             |

These results demonstrate consistent and effective acceleration across various tasks, including coding, math, and general conversation.

## Training Details

- **Training Framework:** This model was trained using **[SpecForge](https://github.com/sgl-project/SpecForge)**, an open-source framework for speculative decoding research.
- **Training Data:** The model was trained on the **EagleChat** dataset. Available on [Hugging Face](https://huggingface.co/datasets/zhaode/EagleChat) and [ModelScope](https://modelscope.cn/datasets/zhaode/EagleChat).
- **Training Duration:** The model was trained for 3 epochs on 8x MI308X GPUs, which took 74 hours and totaled 592 `MI308X GPU-hours`.

</div>

---

<!-- 中文版 README -->
<div id="-zh-cn">

# 适用于 Qwen3-VL-8B-Instruct 的 EAGLE-3 草稿模型

## 模型简介

本仓库包含一个 **EAGLE-3 风格的草稿模型**，专为加速 `Qwen3-VL-8B-Instruct` 大语言模型的推理而训练。

请注意：这是一个**非独立模型**。它必须与对应的基座模型 (`Qwen3-VL-8B-Instruct`) 在推测解码 (speculative decoding) 框架下配合使用，才能实现显著的文本生成加速效果。

- **基座模型:** `Qwen3-VL-8B-Instruct`
- **模型架构:** EAGLE-3 (推测解码草稿模型)
- **核心优势:** 在不牺牲基座模型生成质量的前提下，将文本生成吞吐量提升 1.5 到 2.5 倍。

## 什么是 EAGLE？

EAGLE (Extrapolative A* Generative Language Engine) 是一种先进的推测解码方法。它利用一个轻量的草稿模型并行生成一系列草稿词元 (draft tokens)，然后由更大、更强的基座模型通过单次前向传播进行验证。如果草稿被接受，生成过程就能一次性前进多个步骤，从而实现显著的速度提升。

## 性能表现

本模型在一系列多样化的评测基准上进行了评估。`acc_length` (平均接受的草稿词元数) 反映了加速的效率，数值越高越好。

| 评测基准 (Benchmark) | `acc_length` (num_draft_tokens=4) | `acc_length` (num_draft_tokens=8) |
| :------------------ | :-------------------------------: | :-------------------------------: |
| math500             |               2.30                |               2.55                |
| humaneval           |               2.23                |               2.43                |
| gsm8k               |               2.05                |               2.15                |
| ceval               |               1.84                |               1.93                |
| cmmlu               |               1.81                |               1.88                |
| mtbench             |               1.76                |               1.83                |
| **平均值**           |             **~2.00**             |             **~2.13**             |


这些结果表明，该模型在编码、数学和通用对话等不同任务上都能提供稳定且高效的加速效果。

## 训练细节

- **训练框架:** 本模型使用开源推测解码研究框架 **[SpecForge](https://github.com/sgl-project/SpecForge)** 进行训练。
- **训练数据:** 训练数据使用了 **EagleChat** 数据集。您可以在 [Hugging Face](https://huggingface.co/datasets/zhaode/EagleChat) 或 [ModelScope](https://modelscope.cn/datasets/zhaode/EagleChat) 上获取该数据集。
- **训练耗时:** 训练使用 8x MI308X 训练 3 轮，耗时 74 小时，共 592 `MI308X 卡时`。

</div>
