---
title: qwen2.5-72b-instruct-gptq-int3
canonical_url: "https://www.modelscope.cn/models/tclf90/qwen2.5-72b-instruct-gptq-int3"
md_url: "https://www.modelscope.cn/models/tclf90/qwen2.5-72b-instruct-gptq-int3.md"
repository: tclf90/qwen2.5-72b-instruct-gptq-int3
chinese_name: "通义千问2.5-72B-Chat-GPTQ-Int3-量化修复"
last_updated: 2024-11-01
license: qwen
pipeline_tag: text-generation
tasks:
  - text-generation
model_type:
  - qwen2
architectures:
  - Qwen2ForCausalLM
parameters: 9.7B
tensor_type:
  - I32
  - F16
library_name:
  - safetensors
  - pytorch
frameworks:
  - Pytorch
inference_backends:
  - "deploy_task text/emb"
  - "lmdeploy 0.9.1"
  - "lmdeploy_turbomind 0.9.1"
  - "sglang 0.5.2"
  - "vllm 0.9.2"
downloads: 273
stars: 5
tags:
  - qwen2.5
  - gptq
  - int3
  - "量化修复"
  - vLLM
  - sglang
---

# qwen2.5-72b-instruct-gptq-int3

> qwen2.5-72b-instruct-gptq-int3 - tclf90 在 ModelScope 开源的模型。通义千问2.5-72B-Chat-GPTQ-Int3-量化修复 原模型 qwen/Qwen2.5-72B-Instruct

tclf90/qwen2.5-72b-instruct-gptq-int3 是 ModelScope 魔搭社区上的 9.7B 参数text-generation模型，采用 qwen 许可，可用 deploy_task text/emb、lmdeploy 0.9.1、lmdeploy_turbomind 0.9.1 部署。

- **Repository**: tclf90/qwen2.5-72b-instruct-gptq-int3
- **License**: qwen
- **Tasks**: text-generation
- **Parameters**: 9.7B
- **Inference backends**: deploy_task text/emb, lmdeploy 0.9.1, lmdeploy_turbomind 0.9.1, sglang 0.5.2, vllm 0.9.2
- **Tags**: qwen2.5, gptq, int3, 量化修复, vLLM, sglang
- **Downloads**: 273
- **Stars**: 5
- **Last updated**: 2024-11-01

Source: https://www.modelscope.cn/models/tclf90/qwen2.5-72b-instruct-gptq-int3

---

# 通义千问2.5-72B-Chat-GPTQ-Int3-量化修复
原模型 [qwen/Qwen2.5-72B-Instruct](https://www.modelscope.cn/models/qwen/Qwen2.5-72B-Instruct)


### 【模型更新日期】

注：通过`snapshot_download`函数传入`revision=...`来下载指定的`tag`版本

``` 
2024-11-01
1. add group 128、优化测量算法，提升长文作答能力、支持多卡切片 (tag g128)

2024-09-24 tag g32v2
1. 减少长文时吐字重复与消失的情况

2024-09-21 tag g32
1. add group 32
```

### 【模型列表】

| tag     | 文件大小   | 最近更新时间       |
|---------|--------|--------------|
| `g128` | `31GB` | `2024-11-01` |
| `g32v2` | `35GB` | `2024-09-24` |
| `g32`   | `35GB` | `2024-09-21` |

```python
from modelscope import snapshot_download
snapshot_download('tclf90/qwen2.5-72b-instruct-gptq-int3', cache_dir="本地路径", revision='g128')
```

### 【修复内容】

1. 对GPTQ量化的校准做了额外优化；减少模型的 `1.乱吐字`、`2.无限循环`、`3.长文能力丢失`等情况。
2. 有些推理框架的默认`top_k`与`top_p`较大，可以考虑减小对应数值，来获得更合理的模型输出。
3. 根据模型实际情况，可以支持1卡、2卡及4卡的`tensor-parallel-size`启动。


### 【介绍】
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:

- Significantly **more knowledge** and has greatly improved capabilities in **coding** and **mathematics**, thanks to our specialized expert models in these domains.
- Significant improvements in **instruction following**, **generating long texts** (over 8K tokens), **understanding structured data** (e.g, tables), and **generating structured outputs** especially JSON. **More resilient to the diversity of system prompts**, enhancing role-play implementation and condition-setting for chatbots.
- **Long-context Support** up to 128K tokens and can generate up to 8K tokens.
- **Multilingual support** for over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, Arabic, and more.

For more details, please refer to our [blog](https://qwenlm.github.io/blog/qwen2.5/), [GitHub](https://github.com/QwenLM/Qwen2.5), and [Documentation](https://qwen.readthedocs.io/en/latest/).


### 【模型下载】
```python
from modelscope import snapshot_download
snapshot_download('tclf90/模型名', cache_dir="本地路径", revision='...tag...')
```

### 【高并发RESTFul API推理】
方式1：[vllm](https://github.com/vllm-project/vllm)

方式2：[sglang](https://github.com/sgl-project/sglang)

目前推荐使用sglang进行部署，相较于vllm, sglang于A100实测，能有50%～100%的吞吐增益。
