---
title: Qwen2.5-32B-Instruct-MLX-8bit
canonical_url: "https://www.modelscope.cn/models/okwinds/Qwen2.5-32B-Instruct-MLX-8bit"
md_url: "https://www.modelscope.cn/models/okwinds/Qwen2.5-32B-Instruct-MLX-8bit.md"
repository: okwinds/Qwen2.5-32B-Instruct-MLX-8bit
chinese_name: "通义千问Qwen2.5-32B-Instruct-MLX-8bit"
last_updated: 2024-12-01
license: apache-2.0
pipeline_tag: text-generation
tasks:
  - text-generation
model_type:
  - qwen2
architectures:
  - Qwen2ForCausalLM
parameters: 8.7B
tensor_type:
  - U32
  - F16
library_name:
  - mlx
  - safetensors
  - pytorch
frameworks:
  - Pytorch
language:
  - en
  - zh-cn
inference_backends:
  - "deploy_task text/emb"
  - "lmdeploy 0.9.1"
  - "lmdeploy_turbomind 0.9.1"
  - "sglang 0.5.2"
  - "vllm 0.9.2"
downloads: 135
stars: 1
tags:
  - MLX
  - Instruct
  - Chat
---

# Qwen2.5-32B-Instruct-MLX-8bit

> Qwen2.5-32B-Instruct-MLX-8bit - okwinds 在 ModelScope 开源的模型。Qwen2.5-32B-Instruct-MLX-8bit

okwinds/Qwen2.5-32B-Instruct-MLX-8bit 是 ModelScope 魔搭社区上的 8.7B 参数text-generation模型，采用 apache-2.0 许可，可用 deploy_task text/emb、lmdeploy 0.9.1、lmdeploy_turbomind 0.9.1 部署。

- **Repository**: okwinds/Qwen2.5-32B-Instruct-MLX-8bit
- **License**: apache-2.0
- **Tasks**: text-generation
- **Parameters**: 8.7B
- **Inference backends**: deploy_task text/emb, lmdeploy 0.9.1, lmdeploy_turbomind 0.9.1, sglang 0.5.2, vllm 0.9.2
- **Tags**: MLX, Instruct, Chat
- **Downloads**: 135
- **Stars**: 1
- **Last updated**: 2024-12-01

Source: https://www.modelscope.cn/models/okwinds/Qwen2.5-32B-Instruct-MLX-8bit

---

# Qwen2.5-32B-Instruct-MLX-8bit

## 简介

MLX  是 Apple 机器学习研究团队，推出的一款针对 Apple silicon 芯片的机器学习 array framework。

#### MLX 的主要特性

- MLX提供了一个与NumPy紧密一致的Python API。MLX还提供了功能齐全的C++、C和Swift API，这些API与Python API高度相似。MLX拥有更高级别的包，如mlx.nn和mlx.optimizers，它们的API与PyTorch紧密一致，以简化构建更复杂模型的过程。

- 可组合函数变换：MLX支持可组合的函数变换，用于自动微分、自动向量化以及计算图优化。

- Lazy computation：MLX中的计算是惰性的。数组仅在需要时才会被实际创建。

- 动态图构建：MLX中的计算图是动态构建的。改变函数参数的形状不会触发缓慢的编译过程，调试过程简单直观。

- 多设备支持：操作可以在支持的任何设备上运行（目前包括CPU和GPU）。

- 统一内存模型：MLX与其它框架的主要区别在于其统一内存模型。在MLX中，数组存储在共享内存中，这意味着可以在任何支持的设备类型上执行操作，而无需在设备间传输数据。

## Qwen2.5 同系列模型

#### Finetuned Model

- **Model tree** </br>
	&nbsp;Base model - [Qwen/Qwen2.5-32B](https://www.modelscope.cn/models/Qwen/Qwen2.5-32B) </br>
		&nbsp; <span style="font-weight: bold;">╰</span>─  Finetuned - [Qwen/Qwen2.5-32B-Instruct](https://www.modelscope.cn/models/Qwen/Qwen2.5-32B-Instruct) </br>
		&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;<span style="font-weight: bold;"> ├</span>─  Finetuned -  [CompassJudger-1](https://www.modelscope.cn/collections/CompassJudger-1-Int-and-FP8---Mixed-Precision-85d756858db345) </br>
		&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;<span style="font-weight: bold;">└</span>─  Finetuned -  [Rombos-LLM-Qwen2.5](https://www.modelscope.cn/collections/Rombos-LLM-Qwen25-4dc8690ad7f541) </br>

#### 同系量化

| **Model**    | GPTQ Compression & FP8                                                                                                                                                                                                                                                                                                                                                                                                     | GGUF                                                                                         | **MLX** (macOS)                                                                                                                                                                                                                                                              |
| ------------ |----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| *3B params*  | -                                                                                                                                                                                                                                                                                                                                                                                                                          | [3B-GGUF-V3-LOT](https://www.modelscope.cn/models/okwinds/Qwen2.5-3B-Instruct-GGUF-V3-LOT)   | [3B-MLX-4bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-3B-Instruct-MLX-4bit)<br/>[3B-MLX-8bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-3B-Instruct-MLX-8bit)                                                                                                |
| *7B params*  | [7B-Int4-W4A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-7B-Instruct-Int4-W4A16)<br/>[7B-Int8-W8A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-7B-Instruct-Int8-W8A16)<br/>[7B-FP8-A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-7B-Instruct-FP8-A16) <br/>[Coder-7B-Instruct-FP8](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-7B-Instruct-FP8) | [7B-GGUF-V3-LOT](https://www.modelscope.cn/models/okwinds/Qwen2.5-7B-Instruct-GGUF-V3-LOT)   | [7B-MLX-4bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-7B-Instruct-MLX-4bit)<br/>[7B-MLX-8bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-7B-Instruct-MLX-8bit)                                                                                                |
| *Coder 7B params* | [Coder-7B-Int4-W4A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-7B-Instruct-Int4-W4A16)<br/>[Coder-7B-Int8-W8A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-7B-Instruct-Int8-W8A16) <br/>[Coder-7B-FP8-A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-7B-Instruct-FP8-A16)       | [Coder-7B-GGUF-V3-LOT](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-7B-Instruct-GGUF-V3-LOT) | [Coder-7B-MLX-4bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-7B-Instruct-MLX-4bit)<br/>[Coder-7B-MLX-8bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-7B-Instruct-MLX-8bit)   |
| *14B params* | [14B-Int4-W4A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-14B-Instruct-Int4-W4A16)<br/>[14B-Int8-W8A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-14B-Instruct-Int8-W8A16)<br/>[14B-FP8-A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-14B-Instruct-FP8-A16)                                                                                                                                         | [14B-GGUF-V3-LOT](https://www.modelscope.cn/models/okwinds/Qwen2.5-14B-Instruct-GGUF-V3-LOT) | [14B-MLX-4bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-14B-Instruct-MLX-4bit)<br/>[14B-MLX-8bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-14B-Instruct-MLX-8bit)                                                                                            |
| *Coder 14B params* | [Coder-14B-Int4-W4A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-14B-Instruct-Int4-W4A16)<br/>[Coder-14B-Int8-W8A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-14B-Instruct-Int8-W8A16)                                                                                                                                        | [Coder-14B-GGUF-V3-LOT](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-14B-Instruct-GGUF-V3-LOT) | [Coder-14B-MLX-4bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-14B-Instruct-MLX-4bit)<br/>[Coder-14B-MLX-8bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-14B-Instruct-MLX-8bit)  |
| *32B params* | [32B-Int4-W4A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-32B-Instruct-Int4-W4A16)<br/>[32B-Int8-W8A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-32B-Instruct-Int8-W8A16)<br/>[32B-FP8-A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-32B-Instruct-FP8-A16)                                                                                                                                                                                                                                        | [32B-GGUF-V3-LOT](https://www.modelscope.cn/models/okwinds/Qwen2.5-32B-Instruct-GGUF-V3-LOT) | [32B-MLX-2bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-32B-Instruct-MLX-2bit)<br/>[32B-MLX-4bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-32B-Instruct-MLX-4bit)<br/>[32B-MLX-8bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-32B-Instruct-MLX-8bit) |
| *Coder 32B params* | [Coder-32B-Int4-W4A16](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-32B-Instruct-Int4-W4A16)<br/>[Coder-32B-FP8](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-32B-Instruct-FP8)                                                                                                                                                                                                                                        | [Coder-32B-GGUF-V3-LOT](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-32B-Instruct-GGUF-V3-LOT) | [Coder-32B-MLX-4bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-32B-Instruct-MLX-4bit)<br/>[Coder-32B-MLX-8bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-Coder-32B-Instruct-MLX-8bit) |
| *QwQ-Preview 32B params* | [QwQ-32B-Preview-Int4-W4A16](https://www.modelscope.cn/models/okwinds/QwQ-32B-Preview-Int4-W4A16)<br/>[QwQ-32B-Preview-FP8](https://www.modelscope.cn/models/okwinds/QwQ-32B-Preview-FP8)    | [QwQ-32B-GGUF-V3-LOT](https://www.modelscope.cn/models/okwinds/QwQ-32B-Preview-GGUF-V3-LOT) | [QwQ-32B-MLX-4bit](https://www.modelscope.cn/models/okwinds/QwQ-32B-Preview-MLX-4bit)<br/>[QwQ-32B-MLX-8bit](https://www.modelscope.cn/models/okwinds/QwQ-32B-Preview-MLX-8bit) |
| *72B params* | -                                                                                                                                                                                                                                                                                                                                                                                                                          | [72B-GGUF-V3-LOT](https://www.modelscope.cn/models/okwinds/Qwen2.5-72B-Instruct-GGUF-V3-LOT) | [72B-MLX-2bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-72B-Instruct-MLX-2bit)<br/>[72B-MLX-4bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-72B-Instruct-MLX-4bit)<br/>[72B-MLX-8bit](https://www.modelscope.cn/models/okwinds/Qwen2.5-72B-Instruct-MLX-8bit) |

## 模型概述

Qwen2.5-32B-Instruct-MLX-8bit 是一个基于 通义千问 Qwen2.5-32B-Instruct 的 8bit 量化的模型。

- **模型名称:** Qwen2.5-32B-Instruct-MLX-8bit
- **模型架构:** Qwen2.5
- **权重量化:** 8bit

该模型通过将 [qwen/Qwen2.5-32B-Instruct](https://www.modelscope.cn/models/qwen/Qwen2.5-32B-Instruct) 的 16bit 权重量化为 8bit 而实现。

量化过程将每个参数从 16bit 减少到 8bit，将模型占用磁盘空间大小，以及推理时加载模型需要的 MAC 内存空间，减少到了大约为原模型的1/2。

> <span style="color: red;">注意：MLX 仅适用于运行 macOS >= 13.5 的设备，强烈建议使用 macOS 14 （Sonoma）</span>

## 推理要求

运行以下命令以安装所需的 MLX 软件包。

```
pip install mlx-lm mlx -U
```

## Quickstart

#### 代码示例

这里提供了一个带有 `apply_chat_template` 的代码片段，展示如何加载 tokenizer 和模型以及如何生成内容。

```python
from mlx_lm import load, generate
from modelscope import snapshot_download
model_dir =snapshot_download('okwinds/Qwen2.5-32B-Instruct-MLX-8bit')
model, tokenizer = load(model_dir, tokenizer_config={"eos_token": "<|im_end|>"})

prompt = "天空为什么是蓝色的？"
messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True
)

response = generate(model, tokenizer, prompt=text, verbose=True, top_p=0.8, temp=0.7, repetition_penalty=1.05, max_tokens=512)
```

推理结果样例：

```plaintext
==========
Prompt: <|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
天空为什么是蓝色的？<|im_end|>
<|im_start|>assistant

天空之所以呈现蓝色，是因为地球大气中的气体分子和其他微小颗粒会散射太阳光。这种现象被称为瑞利散射（Rayleigh scattering），是由英国物理学家约翰·威廉·斯特拉特（Lord Rayleigh）在19世纪末提出的。

太阳光是由多种颜色的光组成的，这些颜色混合在一起形成了我们所看到的白光。不同颜色的光波长不同：蓝色光的波长较短，而红色光的波长较长。当太阳光穿过地球大气层时，气体分子和其他微小颗粒会散射这些光线。由于蓝色光的波长较短，因此更容易被气体分子散射到各个方向。这就是为什么我们在白天看到的天空是蓝色的原因。

当太阳接近地平线时，阳光需要穿过更厚的大气层才能到达我们的眼睛，这时更多的蓝色和绿色光线被散射掉，剩下的红色和橙色光线较多，所以太阳看起来是红色或橙色的。
==========
Prompt: 25 tokens, 20.174 tokens-per-sec
Generation: 204 tokens, 10.771 tokens-per-sec
Peak memory: 17.398 GB
```

#### 推理引擎

推荐在 MAC 上使用 xinference 进行推理，xinference 是一个一站式推理 LLM 的引擎，可以方便易用地加载和运行 MLX 模型。并且兼容 Openai 样式的 API，对 Qwen 系列模型支持 Function calling 推理。

使用 conda 创建一个虚拟环境，并安装 xinference ：

```bash
# 创建 python 3.10 虚拟环境
conda create -n xinference python=3.10 -y

# 激活虚拟环境
conda activate xinference

# 安装 xinference 以及 mlx 依赖
pip install "xinference[mlx]"
```

启动 xinference：

```bash
xinference-local -H 0.0.0.0 -p 9997
```

启动后，可以通过http://localhost:9997 访问 xinference 的 web 界面。通过 Register Model 注册下载好的模型以后，即可推理。
