---
title: Mistral-7B-Instruct-v0.2-GGUF
canonical_url: "https://www.modelscope.cn/models/Xorbits/Mistral-7B-Instruct-v0.2-GGUF"
md_url: "https://www.modelscope.cn/models/Xorbits/Mistral-7B-Instruct-v0.2-GGUF.md"
repository: Xorbits/Mistral-7B-Instruct-v0.2-GGUF
last_updated: 2023-12-22
license: "Apache License 2.0"
pipeline_tag: text-generation
tasks:
  - text-generation
library_name:
  - gguf
  - xinference
frameworks:
  - xinference
downloads: 550
stars: 2
tags:
  - gguf
---

# Mistral-7B-Instruct-v0.2-GGUF

> Mistral-7B-Instruct-v0.2-GGUF - Xorbits 在 ModelScope 开源的模型。Mistral-7B-Instruct-v0.2-GGUF

Xorbits/Mistral-7B-Instruct-v0.2-GGUF 是 ModelScope 魔搭社区上的text-generation模型，采用 Apache License 2.0 许可。

- **Repository**: Xorbits/Mistral-7B-Instruct-v0.2-GGUF
- **License**: Apache License 2.0
- **Tasks**: text-generation
- **Tags**: gguf
- **Downloads**: 550
- **Stars**: 2
- **Last updated**: 2023-12-22

Source: https://www.modelscope.cn/models/Xorbits/Mistral-7B-Instruct-v0.2-GGUF

---

## Mistral-7B-Instruct-v0.2-GGUF

This repo contains GGUF format model files for Mistral-7B-Instruct-v0.2.

### About GGUF
GGUF is a new format introduced by the [llama.cpp](https://github.com/ggerganov/llama.cpp) team on August 21st 2023. 
It is a replacement for GGML, which is no longer supported by [llama.cpp](https://github.com/ggerganov/llama.cpp).

Supported quantization methods:
- Q4_K_M

Will add more methods in the future, you can contact us if support for other quantification is needed. 

### Example code

#### Install packages
```bash
pip install -U xinference[ggml]
```
If you want to run with GPU acceleration, refer to [installation](https://github.com/xorbitsai/inference#installation).

####  Start a local instance of Xinference
```bash
xinference -p 9997
```

#### Launch and inference
```python
from xinference.client import Client

client = Client("http://localhost:9997")
model_uid = client.launch_model(
    model_name="mistral-instruct-v0.2",
    model_format="ggufv2", 
    model_size_in_billions=7,
    quantization="Q4_K_M",
    )
model = client.get_model(model_uid)

chat_history = []
prompt = "What is the largest animal?"
model.chat(
    prompt,
    chat_history=chat_history,
    generate_config={"max_tokens": 1024}
)
```

### More information

[Xinference](https://github.com/xorbitsai/inference) Replace OpenAI GPT with another LLM in your app 
by changing a single line of code. Xinference gives you the freedom to use any LLM you need. 
With Xinference, you are empowered to run inference with any open-source language models, 
speech recognition models, and multimodal models, whether in the cloud, on-premises, or even on your laptop.

<i><a href="https://join.slack.com/t/xorbitsio/shared_invite/zt-1z3zsm9ep-87yI9YZ_B79HLB2ccTq4WA">👉 Join our Slack community!</a></i>
