---
title: DeepSeek-V3.2-w4a8
canonical_url: "https://www.modelscope.cn/models/taoxiaoxin/DeepSeek-V3.2-w4a8"
md_url: "https://www.modelscope.cn/models/taoxiaoxin/DeepSeek-V3.2-w4a8.md"
repository: taoxiaoxin/DeepSeek-V3.2-w4a8
last_updated: 2025-12-26
license: mit
pipeline_tag: text-generation
tasks:
  - text-generation
model_type:
  - deepseek_v32
architectures:
  - DeepseekV32ForCausalLM
base_model:
  - deepseek-ai/DeepSeek-V3.2
base_model_relation: quantized
parameters: 360.5B
tensor_type:
  - F32
  - I8
  - BF16
library_name:
  - safetensors
  - pytorch
frameworks:
  - PyTorch
downloads: 447
stars: 4
---

# DeepSeek-V3.2-w4a8

> DeepSeek-V3.2-w4a8 - taoxiaoxin 在 ModelScope 开源的模型。说明：该模型是通过参考msmodelslim 量化而来，参考教程：DeepSeek 量化案例

taoxiaoxin/DeepSeek-V3.2-w4a8 是 ModelScope 魔搭社区上的 360.5B 参数text-generation模型，采用 mit 许可，基于 deepseek-ai/DeepSeek-V3.2 构建。

- **Repository**: taoxiaoxin/DeepSeek-V3.2-w4a8
- **License**: mit
- **Tasks**: text-generation
- **Parameters**: 360.5B
- **Base model**: deepseek-ai/DeepSeek-V3.2
- **Downloads**: 447
- **Stars**: 4
- **Last updated**: 2025-12-26

Source: https://www.modelscope.cn/models/taoxiaoxin/DeepSeek-V3.2-w4a8

---

# DeepSeek-V3.2-w4a8

> 说明：该模型是通过参考`msmodelslim` 量化而来，参考教程：[DeepSeek 量化案例](https://gitcode.com/Ascend/msit/tree/master/msmodelslim/example/DeepSeek#deepseek-v32-exp%E5%90%ABmtp%E5%B1%82-w8a8-%E6%B7%B7%E5%90%88%E9%87%8F%E5%8C%96)

### 基础模型
+ 基础模型: [【deepseek-ai/DeepSeek-V3.2】](https://modelscope.cn/models/deepseek-ai/DeepSeek-V3.2)

### 量化过程
~~~bash
source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:False
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7


# 创建虚拟环境
conda create -y -n msit_env python=3.10
conda activate msit_env

# 安装依赖包
pip3 install attrs cython 'numpy>=1.19.2,<=1.24.0' decorator sympy cffi pyyaml pathlib2 psutil protobuf==3.20.0 scipy requests absl-py --user
pip install torch==7.1.0 torch_npu==2.7.1rc1
pip install transformers==4.48.2
git clone https://gitcode.com/Ascend/msit.git
cd ./msit/msmodelslim
bash install.sh 

# DeepSeek-V3.2(含MTP层) W4A8 混合量化
nohup msmodelslim quant \
 --model_path ./model/DeepSeek-V3.2 \
 --save_path ./model/DeepSeek-V3.2-w4a8 \
 --model_type DeepSeek-V3.2-Exp \
 --quant_type w4a8 \
 --device npu \
 --trust_remote_code True > quant_v3.2_w4a8.log 2>&1 &

tail -f quant_w4a8.log
~~~

### 推理部署(`910B A2 64G*8*2`)

#### 1. 启动docker  容器并进入(两个节点同时执行)

```
export IMAGE=quay.io/ascend/vllm-ascend:v0.12.0rc1
docker run -d --name vllm-ascend \
    --shm-size=1g \
    --net=host \
    --device /dev/davinci0 \
    --device /dev/davinci1 \
    --device /dev/davinci2 \
    --device /dev/davinci3 \
    --device /dev/davinci4 \
    --device /dev/davinci5 \
    --device /dev/davinci6 \
    --device /dev/davinci7 \
    --device /dev/davinci_manager \
    --device /dev/devmm_svm \
    --device /dev/hisi_hdc \
    -v /usr/local/dcmi:/usr/local/dcmi \
    -v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
    -v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
    -v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
    -v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
    -v /etc/ascend_install.info:/etc/ascend_install.info \
    -v /your_model_path:/mnt/data/models \ # 修改为你的模型路径
    -it $IMAGE bash
    
docker exec -it vllm-ascend bash
```

#### 2. node0 设置环境变量

```
cat > ray_node0.sh << 'EOF'
#!/bin/bash
nic_name="你的网卡名称,如:eno0"
local_ip="你的本机内网IP"

source /usr/local/Ascend/ascend-toolkit/set_env.sh
export HCCL_IF_IP=$local_ip
export GLOO_SOCKET_IFNAME=$nic_name
export TP_SOCKET_IFNAME=$nic_name
export HCCL_SOCKET_IFNAME=$nic_name
export RAY_EXPERIMENTAL_NOSET_ASCEND_RT_VISIBLE_DEVICES=1
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
export OMP_PROC_BIND=false
export OMP_NUM_THREADS=1
export HCCL_BUFFSIZE=1024
export VLLM_ENGINE_INIT_TIMEOUT=3600
export ENGINE_INIT_TIMEOUT=3600
export HCCL_EXEC_TIMEOUT=7200
export HCCL_CONNECT_TIMEOUT=3600
export VLLM_USE_V1=1
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
export VLLM_ASCEND_ENABLE_MLAPO=1
EOF
```

#### 3. node1  设置环境变量
```
cat > ray_node1.sh << 'EOF'
#!/bin/bash
nic_name="你的网卡名称,如:eno0"
local_ip="node1 的内网IP"
node0_ip="node0 的内网IP"

source /usr/local/Ascend/ascend-toolkit/set_env.sh
export HCCL_IF_IP=$local_ip
export GLOO_SOCKET_IFNAME=$nic_name
export TP_SOCKET_IFNAME=$nic_name
export HCCL_SOCKET_IFNAME=$nic_name
export RAY_EXPERIMENTAL_NOSET_ASCEND_RT_VISIBLE_DEVICES=1
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
export OMP_PROC_BIND=false
export OMP_NUM_THREADS=1
export HCCL_BUFFSIZE=1024
export VLLM_ENGINE_INIT_TIMEOUT=3600
export ENGINE_INIT_TIMEOUT=3600
export HCCL_EXEC_TIMEOUT=7200
export HCCL_CONNECT_TIMEOUT=3600
export VLLM_USE_V1=1
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
export VLLM_ASCEND_ENABLE_MLAPO=1
EOF
```

#### 4. node0 起 head
```
# node0
source ray_node0.sh

ray stop -f || true

# 显式告诉 Ray：本节点有 8 个 NPU
ray start --head \
  --node-ip-address=10.13.200.15 \
  --port=6379 \
  --resources='{"NPU": 8}'
```

#### 5.node1 起 worker
```
# node1
source ray_node1.sh

ray stop -f || true

ray start --address='10.13.200.15:6379' \
  --node-ip-address=10.13.200.16 \
  --resources='{"NPU": 8}'
```
##### 6.回到 node0 验证资源
```
ray status
python3 - <<'PY'
import ray
ray.init(address="auto")
print("cluster_resources=", ray.cluster_resources())
print("available_resources=", ray.available_resources())
PY

```

#### 7. 启动vllm server 
```
vllm serve /mnt/data/models/DeepSeek-V3.2-w4a8/ \
  --max-model-len 32768 \
  --max-num-batched-tokens 4096 \
  --quantization ascend \
  --enable-expert-parallel \
  --port 8899 \
  --served-model-name DeepSeek-Plus \
  --gpu_memory_utilization 0.92 \
  --max-num-seqs 8 \
  --tensor-parallel-size 8 \
  --pipeline-parallel-size 2 \
  --no-enable-prefix-caching \
  --distributed_executor_backend ray
```

### 【Logs】
```
2025-12-17
1. initial commit
```

### 【Model Files】
| File Size | Last Updated |
|-----------|--------------|
| `350GB`   | `2025-12-17` |

### 【Model Download】
```python
nohup modelscope download --model taoxiaoxin/DeepSeek-V3.2-w4a8 --local_dir your_local_path > dd.log 2>&1 &

tail -f dd.log
```
# DeepSeek-V3.2: Efficient Reasoning & Agentic AI

<!-- markdownlint-disable first-line-h1 -->
<!-- markdownlint-disable html -->
<!-- markdownlint-disable no-duplicate-header -->

<div align="center">
  <img src="https://github.com/deepseek-ai/DeepSeek-V2/blob/main/figures/logo.svg?raw=true" width="60%" alt="DeepSeek-V3" />
</div>
<hr>
<div align="center" style="line-height: 1;">
  <a href="https://www.deepseek.com/" target="_blank" style="margin: 2px;">
    <img alt="Homepage" src="https://github.com/deepseek-ai/DeepSeek-V2/blob/main/figures/badge.svg?raw=true" style="display: inline-block; vertical-align: middle;"/>
  </a>
  <a href="https://chat.deepseek.com/" target="_blank" style="margin: 2px;">
    <img alt="Chat" src="https://img.shields.io/badge/🤖%20Chat-DeepSeek%20V3-536af5?color=536af5&logoColor=white" style="display: inline-block; vertical-align: middle;"/>
  </a>
  <a href="https://huggingface.co/deepseek-ai" target="_blank" style="margin: 2px;">
    <img alt="Hugging Face" src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-DeepSeek%20AI-ffc107?color=ffc107&logoColor=white" style="display: inline-block; vertical-align: middle;"/>
  </a>
</div>
<div align="center" style="line-height: 1;">
  <a href="https://discord.gg/Tc7c45Zzu5" target="_blank" style="margin: 2px;">
    <img alt="Discord" src="https://img.shields.io/badge/Discord-DeepSeek%20AI-7289da?logo=discord&logoColor=white&color=7289da" style="display: inline-block; vertical-align: middle;"/>
  </a>
  <a href="https://github.com/deepseek-ai/DeepSeek-V2/blob/main/figures/qr.jpeg?raw=true" target="_blank" style="margin: 2px;">
    <img alt="Wechat" src="https://img.shields.io/badge/WeChat-DeepSeek%20AI-brightgreen?logo=wechat&logoColor=white" style="display: inline-block; vertical-align: middle;"/>
  </a>
  <a href="https://twitter.com/deepseek_ai" target="_blank" style="margin: 2px;">
    <img alt="Twitter Follow" src="https://img.shields.io/badge/Twitter-deepseek_ai-white?logo=x&logoColor=white" style="display: inline-block; vertical-align: middle;"/>
  </a>
</div>
<div align="center" style="line-height: 1;">
  <a href="LICENSE" style="margin: 2px;">
    <img alt="License" src="https://img.shields.io/badge/License-MIT-f5de53?&color=f5de53" style="display: inline-block; vertical-align: middle;"/>
  </a>
</div>

<p align="center">
  <a href="assets/paper.pdf"><b>Technical Report</b>👁️</a>
</p>

## Introduction

We introduce **DeepSeek-V3.2**, a model that harmonizes high computational efficiency with superior reasoning and agent performance. Our approach is built upon three key technical breakthroughs:

1. **DeepSeek Sparse Attention (DSA):** We introduce DSA, an efficient attention mechanism that substantially reduces computational complexity while preserving model performance, specifically optimized for long-context scenarios.
2. **Scalable Reinforcement Learning Framework:** By implementing a robust RL protocol and scaling post-training compute, *DeepSeek-V3.2* performs comparably to GPT-5. Notably, our high-compute variant, **DeepSeek-V3.2-Speciale**, **surpasses GPT-5** and exhibits reasoning proficiency on par with Gemini-3.0-Pro.
    - *Achievement:* 🥇 **Gold-medal performance** in the 2025 International Mathematical Olympiad (IMO) and International Olympiad in Informatics (IOI).
3. **Large-Scale Agentic Task Synthesis Pipeline:** To integrate **reasoning into tool-use** scenarios, we developed a novel synthesis pipeline that systematically generates training data at scale. This facilitates scalable agentic post-training, improving compliance and generalization in complex interactive environments.

<div align="center">
 <img src="assets/benchmark.png" >
</div>

We have also released the final submissions for IOI 2025, ICPC World Finals, IMO 2025 and CMO 2025, which were selected based on our designed pipeline. These materials are provided for the community to conduct secondary verification. The files can be accessed at `assets/olympiad_cases`.

## Chat Template

DeepSeek-V3.2 introduces significant updates to its chat template compared to prior versions. The primary changes involve a revised format for tool calling and the introduction of a "thinking with tools" capability.

To assist the community in understanding and adapting to this new template, we have provided a dedicated `encoding` folder, which contains Python scripts and test cases demonstrating how to encode messages in OpenAI-compatible format into input strings for the model and how to parse the model's text output.

A brief example is illustrated below:

```python
import transformers
# encoding/encoding_dsv32.py
from encoding_dsv32 import encode_messages, parse_message_from_completion_text

tokenizer = transformers.AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V3.2")

messages = [
    {"role": "user", "content": "hello"},
    {"role": "assistant", "content": "Hello! I am DeepSeek.", "reasoning_content": "thinking..."},
    {"role": "user", "content": "1+1=?"}
]
encode_config = dict(thinking_mode="thinking", drop_thinking=True, add_default_bos_token=True)

# messages -> string
prompt = encode_messages(messages, **encode_config)
# Output: "<｜begin▁of▁sentence｜><｜User｜>hello<｜Assistant｜></think>Hello! I am DeepSeek.<｜end▁of▁sentence｜><｜User｜>1+1=?<｜Assistant｜><think>"

# string -> tokens
tokens = tokenizer.encode(prompt)
# Output: [0, 128803, 33310, 128804, 128799, 19923, 3, 342, 1030, 22651, 4374, 1465, 16, 1, 128803, 19, 13, 19, 127252, 128804, 128798]
```

Important Notes:

1. This release does not include a Jinja-format chat template. Please refer to the Python code mentioned above.
2. The output parsing function included in the code is designed to handle well-formatted strings only. It does not attempt to correct or recover from malformed output that the model might occasionally generate. It is not suitable for production use without robust error handling.
3. A new role named `developer` has been introduced in the chat template. This role is dedicated exclusively to search agent scenarios and is designated for no other tasks. The official API does not accept messages assigned to `developer`.

## How to Run Locally

The model structure of DeepSeek-V3.2 and DeepSeek-V3.2-Speciale are the same as DeepSeek-V3.2-Exp. Please visit [DeepSeek-V3.2-Exp](https://github.com/deepseek-ai/DeepSeek-V3.2-Exp) repo for more information about running this model locally.

Usage Recommendations:

1. For local deployment, we recommend setting the sampling parameters to `temperature = 1.0, top_p = 0.95`.
2. Please note that the DeepSeek-V3.2-Speciale variant is designed exclusively for deep reasoning tasks and does not support the tool-calling functionality.

## License

This repository and the model weights are licensed under the [MIT License](LICENSE).

## Citation

```
@misc{deepseekai2025deepseekv32,
      title={DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models}, 
      author={DeepSeek-AI},
      year={2025},
}
```

## Contact

If you have any questions, please raise an issue or contact us at [service@deepseek.com](service@deepseek.com).
