---
title: Qwen3.5-0.8B-Q4_K_M-GGUF
canonical_url: "https://www.modelscope.cn/models/diodel/Qwen3.5-0.8B-Q4_K_M-GGUF"
md_url: "https://www.modelscope.cn/models/diodel/Qwen3.5-0.8B-Q4_K_M-GGUF.md"
repository: diodel/Qwen3.5-0.8B-Q4_K_M-GGUF
last_updated: 2026-03-14
license: "Apache License 2.0"
pipeline_tag: image-text-to-text
tasks:
  - image-text-to-text
base_model:
  - Qwen/Qwen3.5-0.8B
base_model_relation: quantized
library_name:
  - gguf
  - pytorch
frameworks:
  - Pytorch
downloads: 1083
stars: 1
tags:
  - Qwen3.5
  - qwen3_5
  - gguf
---

# Qwen3.5-0.8B-Q4_K_M-GGUF

> Qwen3.5-0.8B-Q4_K_M-GGUF - diodel 在 ModelScope 开源的模型。Quantization This model is created by quantizing Qwen/Qwen3.5-0.8B to Q4KM using converthftogguf.py and llama-quantize from llama.cpp.

diodel/Qwen3.5-0.8B-Q4_K_M-GGUF 是 ModelScope 魔搭社区上的image-text-to-text模型，采用 Apache License 2.0 许可，基于 Qwen/Qwen3.5-0.8B 构建。

- **Repository**: diodel/Qwen3.5-0.8B-Q4_K_M-GGUF
- **License**: Apache License 2.0
- **Tasks**: image-text-to-text
- **Base model**: Qwen/Qwen3.5-0.8B
- **Tags**: Qwen3.5, qwen3_5, gguf
- **Downloads**: 1083
- **Stars**: 1
- **Last updated**: 2026-03-14

Source: https://www.modelscope.cn/models/diodel/Qwen3.5-0.8B-Q4_K_M-GGUF

---

### Quantization
#### This model is created by quantizing Qwen/Qwen3.5-0.8B to Q4_K_M using convert_hf_to_gguf.py and llama-quantize from llama.cpp.

### Target Users
#### This model supports CPU and heterogeneous (CPU+GPU) inference deployment, making it suitable for users without a GPU.

### Usage Steps
#### 1. Download
SDK Download
```bash
# Install ModelScope
pip install modelscope
```
```python
# Download the model with SDK
from modelscope import snapshot_download
model_dir = snapshot_download('diodel/Qwen3.5-0.8B-Q4_K_M-GGUF')
```
Git Download
```bash
# Download the model with Git
git clone https://www.modelscope.cn/diodel/Qwen3.5-0.8B-Q4_K_M-GGUF.git
```
For more download methods, please refer to the official documentation: https://modelscope.cn/docs/models/download
#### 2. Prepare llama.cpp dependencies

#### 3. Build and Start the Service
```
# Build and start the service
./llama.cpp/build/bin/llama-server \
    --model ./Qwen3.5-0.8B-Q4_K_M.gguf \
    --ctx-size 2048 \
    --host 0.0.0.0 \
    --port 8000
```
