---
title: DeepSeek-V4-Flash-w4a8-mtp
canonical_url: "https://www.modelscope.cn/models/gdydems/DeepSeek-V4-Flash-w4a8-mtp"
md_url: "https://www.modelscope.cn/models/gdydems/DeepSeek-V4-Flash-w4a8-mtp.md"
repository: gdydems/DeepSeek-V4-Flash-w4a8-mtp
last_updated: 2026-05-21
license: "MIT License"
pipeline_tag: text-generation
tasks:
  - text-generation
model_type:
  - deepseek_v4
architectures:
  - DeepseekV4ForCausalLM
base_model:
  - deepseek-ai/DeepSeek-V4-Flash
base_model_relation: quantized
library_name:
  - safetensors
  - pytorch
frameworks:
  - PyTorch
downloads: 2067
stars: 0
---

# DeepSeek-V4-Flash-w4a8-mtp

> DeepSeek-V4-Flash-w4a8-mtp - gdydems 在 ModelScope 开源的模型。DeepSeek-V4-Flash-w4a8-mtp

gdydems/DeepSeek-V4-Flash-w4a8-mtp 是 ModelScope 魔搭社区上的text-generation模型，采用 MIT License 许可，基于 deepseek-ai/DeepSeek-V4-Flash 构建。

- **Repository**: gdydems/DeepSeek-V4-Flash-w4a8-mtp
- **License**: MIT License
- **Tasks**: text-generation
- **Base model**: deepseek-ai/DeepSeek-V4-Flash
- **Downloads**: 2067
- **Stars**: 0
- **Last updated**: 2026-05-21

Source: https://www.modelscope.cn/models/gdydems/DeepSeek-V4-Flash-w4a8-mtp

---

# DeepSeek-V4-Flash-w4a8-mtp

## 1. 基本信息
| 项目 | 信息 |
|:------:|:------:|
| 原始模型名  | DeepSeek-V4-Flash |
| 原始模型链接  | [ deepseek-ai/DeepSeek-V4-Flash ](https://www.modelscope.cn/models/deepseek-ai/DeepSeek-V4-Flash) |
| 精度测试机型  | Atlas 800T A2 1台 |
| 精度测试平台  | docker vllm-ascend |
| 版本  | vllm-ascend:v0.13.0rc3 |
| 链接  | quay.m.daocloud.io/ascend/vllm-ascend:v0.13.0rc3 |


## 2 量化脚本：


## 3 精度测试结果

| 模型名 |  量化格式 | 数据集 | 测试精度 % | 官方精度 % | 备注 |
|:------:|:------:|:------:|:------:|:------:|:------:|
| DeepSeek-V4-Flash-w4a8-mtp | w4a8 | gpqa | 72.22 | 71.2 | V4-Flash Non-Think |
| DeepSeek-V4-Flash-w4a8-mtp | w4a8 | mmlupro | 82.41 | 83.0 | V4-Flash Non-Think |
| DeepSeek-V4-Flash-w4a8-mtp | w4a8 | mmlupro | 85.85 | 86.2 | V4-Flash Max |

* 使用ais_bench,其中`Non-Think`模式`max_out_len = 65536`， `Max`模式`max_out_len = 131072`。精度存在波动，建议多次测试。
