---
title: Qwen3-8B-int4-ov
canonical_url: "https://www.modelscope.cn/models/OpenVINO/Qwen3-8B-int4-ov"
md_url: "https://www.modelscope.cn/models/OpenVINO/Qwen3-8B-int4-ov.md"
repository: OpenVINO/Qwen3-8B-int4-ov
last_updated: 2026-02-27
license: apache-2.0
model_type:
  - qwen3
architectures:
  - Qwen3ForCausalLM
base_model:
  - Qwen/Qwen3-8B
base_model_relation: quantized
library_name:
  - transformer
  - openvino
  - pytorch
frameworks:
  - pytorch
inference_backends:
  - "deploy_task text/emb"
  - "lmdeploy_turbomind 0.9.1"
  - "sglang 0.5.2"
  - "vllm 0.9.2"
downloads: 1473
stars: 0
---

# Qwen3-8B-int4-ov

> Qwen3-8B-int4-ov - OpenVINO 在 ModelScope 开源的模型。Qwen3-8B-int4-ov Model creator: Qwen Original model: Qwen3-8B

OpenVINO/Qwen3-8B-int4-ov 是 ModelScope 魔搭社区上的机器学习模型，采用 apache-2.0 许可，基于 Qwen/Qwen3-8B 构建，可用 deploy_task text/emb、lmdeploy_turbomind 0.9.1、sglang 0.5.2 部署。

- **Repository**: OpenVINO/Qwen3-8B-int4-ov
- **License**: apache-2.0
- **Base model**: Qwen/Qwen3-8B
- **Inference backends**: deploy_task text/emb, lmdeploy_turbomind 0.9.1, sglang 0.5.2, vllm 0.9.2
- **Downloads**: 1473
- **Stars**: 0
- **Last updated**: 2026-02-27

Source: https://www.modelscope.cn/models/OpenVINO/Qwen3-8B-int4-ov

---

# Qwen3-8B-int4-ov
 * Model creator: [Qwen](https://huggingface.co/Qwen)
 * Original model: [Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)

## Description
This is [Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) model converted to the [OpenVINO™ IR](https://docs.openvino.ai/format) (Intermediate Representation) format with weights compressed to INT4 by [NNCF](https://github.com/openvinotoolkit/nncf).

## Quantization Parameters

The quantization was performed using `optimum-cli export openvino` with the following parameters:

 * mode: **INT4_ASYM** 
 * ratio: **1.0** 
 * group_size: **128** 
 * scale_estimation: **True** 
 * dataset: **wikitext2** 

For more information on quantization, check the [OpenVINO model optimization guide](https://docs.openvino.ai/weight_compression).

## Compatibility

The provided OpenVINO™ IR model is compatible with:

* OpenVINO version 2026.0.0 and higher
* Optimum Intel 1.27.0 and higher

## Running Model Inference with [Optimum Intel](https://huggingface.co/docs/optimum/intel/index)

1. Install packages required for using [Optimum Intel](https://huggingface.co/docs/optimum/intel/index) integration with the OpenVINO backend:

```
pip install optimum[openvino]
```

2. Run model inference:

```
from transformers import AutoTokenizer
from optimum.intel.openvino import OVModelForCausalLM

model_id = "OpenVINO/qwen3-8b-int4-ov"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = OVModelForCausalLM.from_pretrained(model_id)

inputs = tokenizer("What is OpenVINO?", return_tensors="pt")

outputs = model.generate(**inputs, max_length=200)
text = tokenizer.batch_decode(outputs)[0]
print(text)
```

For more examples and possible optimizations, refer to the [Inference with Optimum Intel](https://docs.openvino.ai/optimum_intel).

## Running Model Inference with [OpenVINO GenAI](https://github.com/openvinotoolkit/openvino.genai)


1. Install packages required for using OpenVINO GenAI:
```
pip install openvino-genai huggingface_hub
```

2. Download model from HuggingFace Hub:
   
```
import huggingface_hub as hf_hub

model_id = "OpenVINO/qwen3-8b-int4-ov"
model_path = "qwen3-8b-int4-ov"

hf_hub.snapshot_download(model_id, local_dir=model_path)

```

3. Run model inference:

```
import openvino_genai as ov_genai

device = "CPU"
pipe = ov_genai.LLMPipeline(model_path, device)
print(pipe.generate("What is OpenVINO?", max_length=200))
```

More GenAI usage examples can be found in OpenVINO GenAI library [docs](https://docs.openvino.ai/genai_inference) and [samples](https://github.com/openvinotoolkit/openvino.genai?tab=readme-ov-file#openvino-genai-samples)

You can find more detaild usage examples in OpenVINO Notebooks:

- [LLM](https://openvinotoolkit.github.io/openvino_notebooks/?search=LLM)
- [RAG text generation](https://openvinotoolkit.github.io/openvino_notebooks/?search=RAG+system&tasks=Text+Generation)

## Running Model with OpenAI client and [OpenVINO Model Server](https://github.com/openvinotoolkit/model_server)

1a. Deploy model on Windows using [binary package](https://docs.openvino.ai/ovms_baremetal):
```
ovms.exe --rest_port 8000 --source_model OpenVINO/Qwen3-8B-int4-ov --model_repository_path models --tool_parser hermes3 --reasoning_parser qwen3 --target_device GPU --cache_size 2 --task text_generation
```
1b. Deploy model in a Docker container:
```
docker run -d --user $(id -u):$(id -g) --rm -p 8000:8000 -v $(pwd)/models:/models --device /dev/dri --group-add=$(stat -c "%g" /dev/dri/render* | head -n 1) openvino/model_server:latest-gpu \
--rest_port 8000 --model_repository_path models --source_model OpenVINO/Qwen3-8B-int4-ov --tool_parser hermes3 --reasoning_parser qwen3 --target_device GPU --cache_size 2 --task text_generation
```
2. Install the client library:

```
pip install openai
```
3. Run the client:
```
from openai import OpenAI
client = OpenAI(
  base_url="http://localhost:8000/v3",
  api_key="unused"
)

stream = client.chat.completions.create(
    model="OpenVINO/Qwen3-8B-int4-ov",
    messages=[{"role": "user", "content": "Hello."}],
    stream=True,
    extra_body={"chat_template_kwargs": {"enable_thinking": False}},
    tools=[],
)
for chunk in stream:
    if chunk.choices[0].delta.content is not None:
        print(chunk.choices[0].delta.content, end="")
```

Also check how to use this model in an agentic flow with function calling, as shown in the [agentic demo](https://docs.openvino.ai/ovms_batching).


## Limitations

Check the original [model card](https://huggingface.co/Qwen/Qwen3-8B) for limitations.

## Legal information

The original model is distributed under [Apache License Version 2.0](https://huggingface.co/Qwen/Qwen3-8B/blob/main/LICENSE) license. More details can be found in [Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B).

## Disclaimer

Intel is committed to respecting human rights and avoiding causing or contributing to adverse impacts on human rights. See [Intel’s Global Human Rights Principles](https://www.intel.com/content/dam/www/central-libraries/us/en/documents/policy-human-rights.pdf). Intel’s products and software are intended only to be used in applications that do not cause or contribute to adverse impacts on human rights.
