---
title: qwen38-27b-nvfp4-vllm-p3-thor04
canonical_url: "https://www.modelscope.cn/models/navyyang/qwen38-27b-nvfp4-vllm-p3-thor04"
md_url: "https://www.modelscope.cn/models/navyyang/qwen38-27b-nvfp4-vllm-p3-thor04.md"
repository: navyyang/qwen38-27b-nvfp4-vllm-p3-thor04
chinese_name: "Qwen3.8-27B NVFP4 混合量化（vLLM-P3 DRIVE Thor 生产权重）"
last_updated: 2026-10-02
license: apache-2.0
model_type:
  - qwen3_next
architectures:
  - Qwen3NextForCausalLM
parameters: 19.9B
tensor_type:
  - U8
  - BF16
  - F8_E4M3
  - F32
library_name:
  - safetensors
downloads: 45
stars: 0
---

# qwen38-27b-nvfp4-vllm-p3-thor04

> qwen38-27b-nvfp4-vllm-p3-thor04 - navyyang 在 ModelScope 开源的模型。Qwen3.8-27B NVFP4 (HF safetensors) — vLLM-P3 production weights (DRIVE Thor thor04)

navyyang/qwen38-27b-nvfp4-vllm-p3-thor04 是 ModelScope 魔搭社区上的 19.9B 参数机器学习模型，采用 apache-2.0 许可。

- **Repository**: navyyang/qwen38-27b-nvfp4-vllm-p3-thor04
- **License**: apache-2.0
- **Parameters**: 19.9B
- **Downloads**: 45
- **Stars**: 0
- **Last updated**: 2026-10-02

Source: https://www.modelscope.cn/models/navyyang/qwen38-27b-nvfp4-vllm-p3-thor04

---

# Qwen3.8-27B NVFP4 (HF safetensors) — vLLM-P3 production weights (DRIVE Thor thor04)

Production weights of our **vLLM** line on NVIDIA DRIVE Thor (Tegra264, sm_101a,
DriveOS 7.0.3) — the exact files serving on the board (port 8080, 262,144 ctx,
2 concurrent, MTP speculative decoding **at depth 1**, 150-question acceptance:
0 empty answers):

| File | Size (bytes) | Role |
|---|---:|---|
| `model.safetensors` | 23,839,051,704 | main weights — compressed-tensors mixed precision (attn/linear_attn FP8 W8A8, MLP NVFP4 W4A4 g16) |
| `model_mtp.safetensors` | 849,400,392 | MTP speculative-decoding draft head |
| `config.json` / `tokenizer*` / `chat_template.jinja` / `generation_config.json` | — | configs as deployed |

⚠️ **MTP depth must stay at 1** (`"num_speculative_tokens": 1`): depth ≥2 crashes
the engine with a GPU illegal write under concurrent load (confirmed 2026-10-02,
three-arm A/B/C test — see KNOWN-ISSUES #2 in the repo below). Depth 1 measured
17.3–17.7 tok/s single-stream with 88.5% draft acceptance.

✅ **OpenAI Function Calling works** (2026-10-02): serve with
`--enable-auto-tool-choice --tool-call-parser qwen3_xml`. **Parser must be
`qwen3_xml`** (Qwen3-native XML tool format) — the `hermes` parser fails to
extract Qwen3 tool calls and leaks raw XML into `message.content`. Requests
without a `tools` field are unaffected. Verified: structured `tool_calls`
returned for both `tool_choice: "auto"` and `"required"`.

Ready to serve — no conversion needed (vLLM `compressed-tensors` layout).
Apache 2.0 (base: Qwen/Qwen3.8-27B; quantization lineage: unsloth/Qwen3.8-27B-NVFP4).
Provenance in `NOTICE.md`. Verify: `sha256sum -c SHA256SUMS.txt`.

Final config & serve command: [vllm-sm101-replica branch](https://github.com/isenlink/thor-fp8-llm/tree/vllm-sm101-replica) → `docs/10-FINAL-CONFIG-20261001.md`.
Cross-compile binaries (torch/vLLM wheels): [thor04-vllm-p3-sm101-deploy](https://www.modelscope.cn/models/navyyang/thor04-vllm-p3-sm101-deploy).
Companion llama.cpp-line GGUF weights (separate stack): [qwen38-27b-gguf-t4ares-thor01](https://www.modelscope.cn/models/navyyang/qwen38-27b-gguf-t4ares-thor01).
