---
title: Meta-Llama-3.1-8B-Instruct-GGUF
canonical_url: "https://www.modelscope.cn/models/LLM-Research/Meta-Llama-3.1-8B-Instruct-GGUF"
md_url: "https://www.modelscope.cn/models/LLM-Research/Meta-Llama-3.1-8B-Instruct-GGUF.md"
repository: LLM-Research/Meta-Llama-3.1-8B-Instruct-GGUF
last_updated: 2024-09-10
pipeline_tag: text-generation
tasks:
  - text-generation
library_name:
  - gguf
  - pytorch
frameworks:
  - Pytorch
downloads: 4229
stars: 17
tags:
  - llama-cpp
  - gguf-my-repo
  - gguf
---

# Meta-Llama-3.1-8B-Instruct-GGUF

> Meta-Llama-3.1-8B-Instruct-GGUF - LLM-Research 在 ModelScope 开源的模型。LLaMa-3.1-8B-Instrcut的GGUF版本。由LMStudio提供。

- **Repository**: LLM-Research/Meta-Llama-3.1-8B-Instruct-GGUF
- **Tasks**: text-generation
- **Tags**: llama-cpp, gguf-my-repo, gguf
- **Downloads**: 4229
- **Stars**: 17
- **Last updated**: 2024-09-10

Source: https://www.modelscope.cn/models/LLM-Research/Meta-Llama-3.1-8B-Instruct-GGUF

---

# leafspark/Meta-Llama-3.1-8B-Instruct-GGUF
This model was converted to GGUF format from [`meta-llama/Meta-Llama-3.1-8B-Instruct`](https://huggingface.co/meta-llama/Meta-Llama-3.1-8B-Instruct) using llama.cpp via the ggml.ai's [GGUF-my-repo](https://huggingface.co/spaces/ggml-org/gguf-my-repo) space.
Refer to the [original model card](https://huggingface.co/meta-llama/Meta-Llama-3.1-8B-Instruct) for more details on the model.

**Quants:**
- q2_k
- q3_k_m

## Use with llama.cpp
Install llama.cpp through brew (works on Mac and Linux)

```bash
brew install llama.cpp

```
Invoke the llama.cpp server or the CLI.

### CLI:
```bash
llama-cli --hf-repo leafspark/Meta-Llama-3.1-8B-Instruct-hf-Q3_K_M-GGUF --hf-file meta-llama-3.1-8b-instruct-hf-q3_k_m.gguf -p "The meaning to life and the universe is"
```

### Server:
```bash
llama-server --hf-repo leafspark/Meta-Llama-3.1-8B-Instruct-hf-Q3_K_M-GGUF --hf-file meta-llama-3.1-8b-instruct-hf-q3_k_m.gguf -c 2048
```

Note: You can also use this checkpoint directly through the [usage steps](https://github.com/ggerganov/llama.cpp?tab=readme-ov-file#usage) listed in the Llama.cpp repo as well.

Step 1: Clone llama.cpp from GitHub.
```
git clone https://github.com/ggerganov/llama.cpp
```

Step 2: Move into the llama.cpp folder and build it with `LLAMA_CURL=1` flag along with other hardware-specific flags (for ex: LLAMA_CUDA=1 for Nvidia GPUs on Linux).
```
cd llama.cpp && LLAMA_CURL=1 make
```

Step 3: Run inference through the main binary.
```
./llama-cli --hf-repo leafspark/Meta-Llama-3.1-8B-Instruct-hf-Q3_K_M-GGUF --hf-file meta-llama-3.1-8b-instruct-hf-q3_k_m.gguf -p "The meaning to life and the universe is"
```
or 
```
./llama-server --hf-repo leafspark/Meta-Llama-3.1-8B-Instruct-hf-Q3_K_M-GGUF --hf-file meta-llama-3.1-8b-instruct-hf-q3_k_m.gguf -c 2048
```
