---
title: llama-cpp-K2-FAST2
canonical_url: "https://www.modelscope.cn/models/Dalvlad/llama-cpp-K2-FAST2"
md_url: "https://www.modelscope.cn/models/Dalvlad/llama-cpp-K2-FAST2.md"
repository: Dalvlad/llama-cpp-K2-FAST2
last_updated: 2026-10-08
license: apache-2.0
base_model:
  - MBZUAI-IFM/llama.cpp
base_model_relation: finetune
downloads: 0
stars: 0
tags:
  - llama.cpp
  - k2-horizon
  - sliding-window-attention
  - long-context
  - fast-inference
---

# llama-cpp-K2-FAST2

> llama-cpp-K2-FAST2 - Dalvlad 在 ModelScope 开源的模型。llama.cpp — K2-FAST2 fork 🚀

Dalvlad/llama-cpp-K2-FAST2 是 ModelScope 魔搭社区上的机器学习模型，采用 apache-2.0 许可，基于 MBZUAI-IFM/llama.cpp 构建。

- **Repository**: Dalvlad/llama-cpp-K2-FAST2
- **License**: apache-2.0
- **Base model**: MBZUAI-IFM/llama.cpp
- **Tags**: llama.cpp, k2-horizon, sliding-window-attention, long-context, fast-inference
- **Downloads**: 0
- **Stars**: 0
- **Last updated**: 2026-10-08

Source: https://www.modelscope.cn/models/Dalvlad/llama-cpp-K2-FAST2

---

# llama.cpp — K2-FAST2 fork 🚀

**This fork makes K2-Horizon GGUFs up to 3.1× faster at long context.**
It is MBZUAI-IFM's `llama.cpp` `model/K2Horizon` branch (commit `42adf01`) plus one addition:
**sliding-window attention (SWA) support for `k2-horizon`** — 51 lines in `src/models/k2-horizon.cpp`.

> 📦 Companion model (self-contained, window embedded): **K2-Horizon-MoVA-36B-A4B-FAST2**
> — [HuggingFace](https://huggingface.co/VladHong/K2-Horizon-MoVA-36B-A4B-FAST2) ·
> [ModelScope](https://modelscope.cn/models/Dalvlad/K2-Horizon-MoVA-36B-A4B-FAST2)

## Measured speed increase vs vanilla APEX-Mini

Same binary, same frozen 61,337-token prompt, same context window, 5 interleaved rounds,
RTX 2080 Ti 22 GB, all layers in VRAM. Full data: `results-head2head.json` in the model repo.

| @ 61K context (61,354-token prompt) | decode tok/s | prefill tok/s | speedup |
|---|---|---|---|
| vanilla APEX-Mini (q8_0 KV — the only way vanilla fits 22 GB at 61K; default f16 KV OOMs) | 16.48 | 692 | — |
| **FAST2 (this fork + embedded W=8192 window)** | **50.72** | **1,482** | **decode 3.1×, prefill 2.1×** |
| FAST2 with W=16384 | 43.90 | 1,187 | decode 2.7× |

At 32K context (W=8192): decode 52.64 vs 36.78 tok/s (**1.43×**), prefill 1,051 vs 949.
At short context FAST2 equals vanilla — decode there is weight-bandwidth-bound and untouched.

**Quality:** perplexity inside the window is numerically identical to full attention
(2.1814 vs 2.1814 on a 240K-token stratified corpus); needle-in-haystack at 61K passes for
facts inside the window. Recall is limited to the last W tokens by design — pick W to match
your workload (W=16384 is a one-line rebuild, see the model repo's `make_swa_gguf.py`).

## What the patch does

K2-Horizon runs **full attention in all 48 layers** (~192 KiB of KV per token — 6+ GiB read
per generated token at 61K context), which is why vanilla decode collapses from ~57 tok/s at
4K to ~16 tok/s at 61K. The patch lets a GGUF declare a sliding window
(`k2-horizon.attention.sliding_window`), which:

1. `load_arch_hparams`: reads the key (from GGUF metadata or `--override-kv`), sets
   `n_swa`, `swa_type = LLAMA_SWA_TYPE_STANDARD`, and marks the attention layers SWA;
2. `graph()`: routes attention through the iswa input path (`build_attn_inp_kv_iswa`) so each
   layer's KV is capped at the last W tokens — decode cost flattens instead of growing.

Without a window key, behavior is bit-identical to stock `model/K2Horizon`.

## Using it

```bash
# build (same as the MBZUAI fork, plus FA_ALL_QUANTS for mixed KV quant support)
cmake -B build -DGGML_CUDA=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DLLAMA_CURL=OFF \
      -DCMAKE_CUDA_ARCHITECTURES=<your-sm> -DCMAKE_BUILD_TYPE=Release
cmake --build build --target llama-server -j

# option A: the FAST2 GGUF (window embedded — no flags)
llama-server -m K2-Horizon-MoVA-36B-A4B-FAST2-SWA8K.gguf -c 65536 -ngl 99 --parallel 1 \
    -fa on -b 2048 -ub 2048 -t 8 --port 8099

# option B: any k2-horizon GGUF + window at load time
llama-server -m K2-Horizon-MoVA-36B-A4B-APEX-Mini.gguf -c 65536 -ngl 99 \
    --override-kv k2-horizon.attention.sliding_window=int:8192 ...
```

## Where to get what

| artifact | HuggingFace | ModelScope |
|---|---|---|
| FAST2 GGUF (window embedded) + card + measurements | [VladHong/K2-Horizon-MoVA-36B-A4B-FAST2](https://huggingface.co/VladHong/K2-Horizon-MoVA-36B-A4B-FAST2) | [Dalvlad/K2-Horizon-MoVA-36B-A4B-FAST2](https://modelscope.cn/models/Dalvlad/K2-Horizon-MoVA-36B-A4B-FAST2) |
| this fork as a source snapshot | [VladHong/llama-cpp-K2-FAST2](https://huggingface.co/VladHong/llama-cpp-K2-FAST2) | branch `k2-fast2` of [Dalvlad/llama-cpp-K2-FAST2](https://modelscope.cn/models/Dalvlad/llama-cpp-K2-FAST2) |
| standalone patch (apply to stock MBZUAI fork) | `k2-fast2-swa.patch` in either model repo | same |

The upstream llama.cpp README is preserved as `README-upstream.md`. All credit for the
`model/K2Horizon` support belongs to MBZUAI-IFM; the FAST2 change is the 51-line diff above.
