---
title: K2-Horizon-MoVA-36B-A4B-FAST2
canonical_url: "https://www.modelscope.cn/models/Dalvlad/K2-Horizon-MoVA-36B-A4B-FAST2"
md_url: "https://www.modelscope.cn/models/Dalvlad/K2-Horizon-MoVA-36B-A4B-FAST2.md"
repository: Dalvlad/K2-Horizon-MoVA-36B-A4B-FAST2
last_updated: 2026-10-08
license: apache-2.0
pipeline_tag: text-generation
tasks:
  - text-generation
base_model:
  - IFM/K2-Horizon-MoVA-36B-A4B
base_model_relation: finetune
language:
  - en
downloads: 2
stars: 0
tags:
  - gguf
  - k2-horizon
  - mova
  - moe
  - long-context
  - fast-inference
  - sliding-window-attention
---

# K2-Horizon-MoVA-36B-A4B-FAST2

> K2-Horizon-MoVA-36B-A4B-FAST2 - Dalvlad 在 ModelScope 开源的模型。K2-Horizon-MoVA-36B-A4B FAST2 🚀

Dalvlad/K2-Horizon-MoVA-36B-A4B-FAST2 是 ModelScope 魔搭社区上的text-generation模型，采用 apache-2.0 许可，基于 IFM/K2-Horizon-MoVA-36B-A4B 构建。

- **Repository**: Dalvlad/K2-Horizon-MoVA-36B-A4B-FAST2
- **License**: apache-2.0
- **Tasks**: text-generation
- **Base model**: IFM/K2-Horizon-MoVA-36B-A4B
- **Tags**: gguf, k2-horizon, mova, moe, long-context, fast-inference, sliding-window-attention
- **Downloads**: 2
- **Stars**: 0
- **Last updated**: 2026-10-08

Source: https://www.modelscope.cn/models/Dalvlad/K2-Horizon-MoVA-36B-A4B-FAST2

---

# K2-Horizon-MoVA-36B-A4B **FAST2** 🚀

### 3.1× faster decode at 61K context. 2.1× faster prefill. Weights **bit-identical** to APEX-Mini. Zero perplexity loss inside the window.

**This is a recipe repo — it deliberately contains no model weights.**
FAST2 is not a different file: it is the original
[Dalvlad/K2-Horizon-MoVA-36B-A4B-APEX-Mini.gguf](https://modelscope.cn/models/Dalvlad/K2-Horizon-MoVA-36B-A4B-APEX-Mini-GGUF)
plus a **160-byte metadata key** (an 8192-token sliding-window attention declaration) and the
companion llama.cpp fork that honors it. Build it yourself in under a minute, below.

## The numbers (measured, single session, 61,337-token prompt, 5 interleaved rounds)

| @ 61K context | vanilla APEX-Mini* | **FAST2** | speedup |
|---|---|---|---|
| **decode tok/s** | 16.48 | **50.72** | **3.07×** |
| **prefill tok/s** | 692 | **1,482** | **2.14×** |

| @ 32K context | vanilla | **FAST2** | speedup |
|---|---|---|---|
| **decode tok/s** | 36.78 | **52.64** | **1.43×** |

At short context FAST2 matches vanilla (~59 tok/s — weight-bandwidth bound, untouched by design).
Raw aggregates and spreads (0.4–2.1%): `results-head2head.json`.

\* Vanilla with its default f16 KV cache **cannot serve 61K context on a 22 GB GPU at all** (OOM);
the comparator is vanilla with q8_0 KV — the only configuration in which vanilla fits.

## Get it running (two steps)

**1. Get the base GGUF** — [Dalvlad/K2-Horizon-MoVA-36B-A4B-APEX-Mini-GGUF](https://modelscope.cn/models/Dalvlad/K2-Horizon-MoVA-36B-A4B-APEX-Mini-GGUF)

**2. Get the patched llama.cpp** — [Dalvlad/llama-cpp-K2-FAST2](https://modelscope.cn/models/Dalvlad/llama-cpp-K2-FAST2) (branch **`k2-fast2`** — full source; the default branch is a pointer card).
Build it exactly like the MBZUAI fork, plus `-DGGML_CUDA_FA_ALL_QUANTS=ON`:

```bash
cmake -B build -DGGML_CUDA=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DLLAMA_CURL=OFF \
      -DCMAKE_CUDA_ARCHITECTURES=<your-sm> -DCMAKE_BUILD_TYPE=Release
cmake --build build --target llama-server -j
```

**Then either (a) window at load time — no file changes at all:**

```bash
llama-server -m K2-Horizon-MoVA-36B-A4B-APEX-Mini.gguf -c 65536 -ngl 99 --parallel 1 \
    -fa on -ctk q4_0 -ctv f16 -b 2048 -ub 2048 -t 8 \
    --override-kv k2-horizon.attention.sliding_window=int:8192 --port 8099
```

**or (b) bake the window into a self-contained GGUF (~12 seconds, from this repo's `make_swa_gguf.py`):**

```bash
python3 make_swa_gguf.py K2-Horizon-MoVA-36B-A4B-APEX-Mini.gguf 8192 K2-FAST2-SWA8K.gguf
llama-server -m K2-FAST2-SWA8K.gguf -c 65536 -ngl 99 --parallel 1 -fa on -b 2048 -ub 2048 -t 8 --port 8099
```

`fast2-serve.sh` wraps all of this with VRAM gating and context-aware flags
(`FAST2_SWA=8192` to enable the window through the launcher).

## Quality: measured, not hand-waved

- **Perplexity inside the window is numerically identical to full attention**
  (2.1814 vs 2.1814 on a 240K-token stratified corpus; Δ 0.0006 at 16K context with W=16384).
- 61K-context needle retrieval: **PASS** inside the window — exact answer; outside it the model
  declines instead of hallucinating (clean failure mode).
- The tradeoff is explicit: **attention recall covers the last W tokens** (8192 here; rebuild with
  W=16384 → 43.90 tok/s @61K, still 2.7× vanilla, recall 16K).

## Why it's fast (one paragraph)

K2-Horizon decodes slowly at long context because all 48 layers run full attention over the
entire history (~192 KiB of KV per token — 6+ GiB read per generated token at 61K). The window
caps each layer's KV, so decode cost flattens instead of growing, and the f16 KV path (measured
23% faster than q8_0 on this kernel despite 2× the bytes) becomes affordable at 61K. We also
measured and rejected the "obvious" tricks: KV quantization below f16 is *slower* here (dequant
cost > bytes saved), and bigger micro-batches only help prefill.

## Files in this repo

| file | what |
|---|---|
| `k2-fast2-swa.patch` | the fork patch (51 lines, applies cleanly to stock MBZUAI `model/K2Horizon`) |
| `make_swa_gguf.py` | bake the window key into any k2-horizon GGUF (choose your own W) |
| `fast2-serve.sh` | tuned launcher (VRAM-gated, ctx-aware, `FAST2_SWA` support) |
| `results-head2head.json` | every measurement behind the claims |
| `checksums.json` | sha256 of the benchmarked artifact + these tools |

## Honesty clause

Every claim comes from gated, interleaved, repeated measurement published in
`results-head2head.json` — the same discipline that showed the *previous* "FAST" build of this
model was not faster at all. FAST2 changes no weights (tensors byte-identical to APEX-Mini,
sha-verified), does not touch short-context speed, and removes attention recall beyond the
window by design. Provenance: base weights IFM K2-Horizon-MoVA-36B-A4B (apache-2.0); APEX-Mini
quant by the reference repo; measured on RTX 2080 Ti 22 GB.
