---
title: Llama-2-13b-Chat-GGUF
canonical_url: "https://www.modelscope.cn/models/Xorbits/Llama-2-13b-Chat-GGUF"
md_url: "https://www.modelscope.cn/models/Xorbits/Llama-2-13b-Chat-GGUF.md"
repository: Xorbits/Llama-2-13b-Chat-GGUF
last_updated: 2023-10-19
license: "Apache License 2.0"
pipeline_tag: text-generation
tasks:
  - text-generation
library_name:
  - gguf
  - xinference
frameworks:
  - xinference
downloads: 358
stars: 4
tags:
  - gguf
---

# Llama-2-13b-Chat-GGUF

> Llama-2-13b-Chat-GGUF - Xorbits 在 ModelScope 开源的模型。Llama-2-13b-Chat-GGUF

Xorbits/Llama-2-13b-Chat-GGUF 是 ModelScope 魔搭社区上的text-generation模型，采用 Apache License 2.0 许可。

- **Repository**: Xorbits/Llama-2-13b-Chat-GGUF
- **License**: Apache License 2.0
- **Tasks**: text-generation
- **Tags**: gguf
- **Downloads**: 358
- **Stars**: 4
- **Last updated**: 2023-10-19

Source: https://www.modelscope.cn/models/Xorbits/Llama-2-13b-Chat-GGUF

---

## Llama-2-13b-Chat-GGUF

This repo contains GGUF format model files for Llama-2-13b-Chat.

### About GGUF
GGUF is a new format introduced by the [llama.cpp](https://github.com/ggerganov/llama.cpp) team on August 21st 2023. 
It is a replacement for GGML, which is no longer supported by [llama.cpp](https://github.com/ggerganov/llama.cpp).

Supported quantization methods:
- Q4_K_M

Will add more methods in the future, you can contact us if support for other quantification is needed. 

### Example code

#### Install packages
```bash
pip install xinference[ggml]>=0.4.3
```
If you want to run with GPU acceleration, refer to [installation](https://github.com/xorbitsai/inference#installation).

####  Start a local instance of Xinference
```bash
xinference -p 9997
```

#### Launch and inference
```python
from xinference.client import Client

client = Client("http://localhost:9997")
model_uid = client.launch_model(
    model_name="llama-2-chat",
    model_format="ggufv2", 
    model_size_in_billions=13,
    quantization="Q4_K_M",
    )
model = client.get_model(model_uid)

chat_history = []
prompt = "What is the largest animal?"
model.chat(
    prompt,
    chat_history=chat_history,
    generate_config={"max_tokens": 1024}
)
```

### More information

[Xinference](https://github.com/xorbitsai/inference) Replace OpenAI GPT with another LLM in your app 
by changing a single line of code. Xinference gives you the freedom to use any LLM you need. 
With Xinference, you are empowered to run inference with any open-source language models, 
speech recognition models, and multimodal models, whether in the cloud, on-premises, or even on your laptop.

<i><a href="https://join.slack.com/t/xorbitsio/shared_invite/zt-1z3zsm9ep-87yI9YZ_B79HLB2ccTq4WA">👉 Join our Slack community!</a></i>
