---
title: altar-1
canonical_url: "https://www.modelscope.cn/models/AikidoSec/altar-1"
md_url: "https://www.modelscope.cn/models/AikidoSec/altar-1.md"
repository: AikidoSec/altar-1
last_updated: 2026-09-23
license: other
pipeline_tag: text-generation
tasks:
  - text-generation
model_type:
  - glm_moe_dsa
architectures:
  - GlmMoeDsaForCausalLM
base_model:
  - cyankiwi/GLM-5.3-AWQ-INT4
base_model_relation: quantized
parameters: 500.8B
tensor_type:
  - I32
  - BF16
  - I64
  - F32
library_name:
  - safetensors
  - pytorch
frameworks:
  - pytorch
downloads: 398
stars: 0
tags:
  - glm
  - glm-5.3
  - moe
  - w4a16
  - awq
  - int4
  - compressed-tensors
  - reap
  - expert-pruning
  - hopper
---

# altar-1

> altar-1 - AikidoSec 在 ModelScope 开源的模型。Altar-1 — a 504B parameter Prune of GLM-5.3

AikidoSec/altar-1 是 ModelScope 魔搭社区上的 500.8B 参数text-generation模型，采用 other 许可，基于 cyankiwi/GLM-5.3-AWQ-INT4 构建。

- **Repository**: AikidoSec/altar-1
- **License**: other
- **Tasks**: text-generation
- **Parameters**: 500.8B
- **Base model**: cyankiwi/GLM-5.3-AWQ-INT4
- **Tags**: glm, glm-5.3, moe, w4a16, awq, int4, compressed-tensors, reap, expert-pruning, hopper
- **Downloads**: 398
- **Stars**: 0
- **Last updated**: 2026-09-23

Source: https://www.modelscope.cn/models/AikidoSec/altar-1

---

# Altar-1 — a 504B parameter Prune of GLM-5.3 

**GLM-5.3 with 34% of its experts removed, at INT4 — 328 GB, built to serve on 4× H200 (Hopper) in vLLM.**
Altar-1 was calibrated on cybersecurity traces, coding, tool calling, reasoning, and English. Additionally we used multi-lingual wikipedia articles.

## What this is

GLM-5.3 is a 753B mixture-of-experts model: each token uses 8 of 256 expert sub-networks per layer (~40B active). **REAP** (Router-weighted Expert Activation Pruning) scores each expert’s real contribution and deletes the least useful ones — no retraining. This cut keeps **168 of 256** experts per layer.

The experts are then INT4 **W4A16** (compressed-tensors, AWQ), taken from the [cyankiwi/GLM-5.3-AWQ-INT4](https://huggingface.co/cyankiwi/GLM-5.3-AWQ-INT4) base. Only the routed experts are 4-bit; attention, the shared expert, the dense layers, and the head stay BF16. vLLM auto-selects the Marlin MoE kernel. Routing is untouched: 8 experts per token out of the 168 that remain, ~40B active parameters, same as the unpruned model.

## How close to the original is it?

**KL divergence vs full BF16: 0.506 nats** (sealed 25-prompt panel, full 154k vocabulary). KL is the standard “how differently do these two models predict” score — **0 = identical**, lower = closer. For reference, an EXL3 build of the same 168-expert cut measures 0.511 — at this bit-width the quantization format barely moves the result. Full numbers: [fidelity study](https://huggingface.co/datasets/0xSero/glm-5.3-reap-fidelity-study).

## Why these experts

Instead of keeping the globally most-frequent experts (which deletes a domain’s specialists), each expert is scored by its **largest share of any single domain’s routed work**, so every domain — code, rare languages, structured output — keeps its specialists. Head-to-head vs frequency pruning: [fidelity study](https://huggingface.co/datasets/0xSero/glm-5.3-reap-fidelity-study).

## Serving (vLLM, 4× H200)

```bash
vllm serve aikido/altar-1 --tensor-parallel-size 4 --trust-remote-code --max-model-len 131072
```

Requires Hopper (H100/H200). 328 GB of weights across 4× H200 leaves room for a 128k-context KV cache at production batch sizes; vLLM selects the Marlin MoE kernel automatically.

## Credits

- **[Z.AI / zai-org](https://huggingface.co/zai-org)** — [GLM-5.3](https://huggingface.co/zai-org/GLM-5.3), the base model.
- **[cyankiwi](https://huggingface.co/cyankiwi)** — the [GLM-5.3-AWQ-INT4](https://huggingface.co/cyankiwi/GLM-5.3-AWQ-INT4) W4A16 base this prune is built on.
- **[Cerebras Research](https://github.com/CerebrasResearch/reap)** — REAP ([arXiv:2510.13999](https://arxiv.org/abs/2510.13999)).
- **[0xSero](https://huggingface.co/0xSero)** — performed the REAP prune, and released the [569B](https://huggingface.co/0xSero/GLM-5.3-569B-W4A16) and [EXL3](https://huggingface.co/0xSero/GLM-5.3-500B-EXL3-3.0bpw) builds of the same cut.

Observations: [`glm-5.3-reap-observations-v1`](https://huggingface.co/datasets/0xSero/glm-5.3-reap-observations-v1) · Fidelity study: [`glm-5.3-reap-fidelity-study`](https://huggingface.co/datasets/0xSero/glm-5.3-reap-fidelity-study) · Built on 8× NVIDIA RTX PRO 6000 Blackwell.

## Deploying Altar

To get help deploying this model to your organization, contact yannick@aikido.dev

If you want to put this model to the test, some of Aikido's products are already powered by Altar, try them today:
- **[Aikido Attack](https://www.aikido.dev/platform/attack)**
- **[AI Code Analysis](https://www.aikido.dev/code/code-audit)**
- **[Deep Review](https://help.aikido.dev/deep-review/how-deep-review-works)**

## License
Inherits the [GLM-5.3 license](https://huggingface.co/zai-org/GLM-5.3).
