---
title: thor01-qwen38-27b-dflash2-t4ares-deploy
canonical_url: "https://www.modelscope.cn/models/navyyang/thor01-qwen38-27b-dflash2-t4ares-deploy"
md_url: "https://www.modelscope.cn/models/navyyang/thor01-qwen38-27b-dflash2-t4ares-deploy.md"
repository: navyyang/thor01-qwen38-27b-dflash2-t4ares-deploy
chinese_name: "Thor01 Qwen3.8-27B DFlash2 262K 部署二进制（T4/ARES 内核）"
last_updated: 2026-10-01
license: apache-2.0
language:
  - en
  - zh
downloads: 0
stars: 0
tags:
  - llama.cpp
  - NVIDIA-Drive-Thor
  - sm_101a
  - NVFP4
  - speculative-decoding
  - dflash
  - automotive
---

# thor01-qwen38-27b-dflash2-t4ares-deploy

> thor01-qwen38-27b-dflash2-t4ares-deploy - navyyang 在 ModelScope 开源的模型。Thor01 Qwen3.8-27B DFlash2 262K 部署二进制（T4/ARES 内核）

navyyang/thor01-qwen38-27b-dflash2-t4ares-deploy 是 ModelScope 魔搭社区上的机器学习模型，采用 apache-2.0 许可。

- **Repository**: navyyang/thor01-qwen38-27b-dflash2-t4ares-deploy
- **License**: apache-2.0
- **Tags**: llama.cpp, NVIDIA-Drive-Thor, sm_101a, NVFP4, speculative-decoding, dflash, automotive
- **Downloads**: 0
- **Stars**: 0
- **Last updated**: 2026-10-01

Source: https://www.modelscope.cn/models/navyyang/thor01-qwen38-27b-dflash2-t4ares-deploy

---

# Thor01 Qwen3.8-27B DFlash2 262K 部署二进制（T4/ARES 内核）

> NVIDIA DRIVE Thor（p3960 / Tegra264 / sm_101a / DriveOS 7.0.3 / CUDA 12.8）上
> llama.cpp 生产部署二进制：Qwen3.8-27B（NVFP4 FFN + F8 注意力，内部量化件）+ DFlash2 投机解码
> + 多模态（mmproj），262144 上下文。200K decode 实测 21.9–22.5 t/s，短提示 34.5–34.8 t/s。

## 文件

| 路径 | 说明 | sha256（前8） |
|---|---|---|
| `bin/llama-server-m2c-t4` | 预编译服务端（ELF aarch64，含 sm_101a cubin，DFlash2+mmproj 支持） | `2066e681…` |
| `bin/deploy.sh` | start/stop/restart/status 管理脚本 | `3641d785…` |
| `src/llama-cpp-thor-t4ares-workingtree-20260928.tar.gz` | 完整工作树快照（推荐源码复现用） | `2ef20ca5…` |
| `src/workingtree-diff.patch` | 相对 llama.cpp `72797e89` 的补丁（104KB，26 文件） | `da830c17…` |
| `src/base-commit.txt` | 上游 pin commit | `c54c6982…` |
| `src/build-thor-cuda128.sh` + `src/toolchain-thor-cuda128-sbsa.cmake` | 交叉编译脚本与工具链 | 见 SHA256SUMS.txt |
| `test/bench8091.py` | 测速脚本（/completion，服务端 timings） | `944c95bc…` |
| `test/*.jsonl` | 三次 200K 实测原始数据 + 短提示实测 | 见 SHA256SUMS.txt |
| `SHA256SUMS.txt` | 全部 13 件校验和 | — |

## 快速开始

```bash
# 1) 校验
sha256sum -c SHA256SUMS.txt
# 2) 板端部署（模型自备，见下方模型清单）
bash bin/deploy.sh start
# 3) 验证
curl --noproxy '*' http://<板IP>:8080/health
# 4) 测速
python3 test/bench8091.py 199980 mytrial
```

完整构建/部署/测试文档见 GitHub 仓库：
[isenlink/thor-fp8-llm](https://github.com/isenlink/thor-fp8-llm)（分支 `thor01-t4ares-deploy`）。

## 模型清单（不随本仓分发）

| 件 | 来源 |
|---|---|
| 主模型 RadixArk-F8attn-v2.gguf（NVFP4 FFN + F8 注意力） | 内部量化件未公开发布（sha `d15dac91…2e7e9`，需自行按上游 [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) 转换或获取） |
| 草稿 Qwen3.8-27B-DFlash2-BF16.gguf | [z-lab/Qwen3.8-27B-DFlash2](https://huggingface.co/z-lab/Qwen3.8-27B-DFlash2) |
| 多模态投影 mmproj（HauhauCS） | HuggingFace 作者 HauhauCS |

## 实测数据（真实硬件）

| 口径 | decode t/s | prefill t/s | 备注 |
|---|---|---|---|
| 200K prompt + 384 out（09-28 冻结验收） | **22.46** | ~133 | draft acceptance 0.51 |
| 200K 生产抽测（09-30） | **21.88** | 134.3 | acc 0.5103 |
| 200K 复测（10-01） | **22.11** | ~133 | mean_len 4.45 |
| 短提示 2K（10-01） | **34.51 / 34.76 / 34.77** | 220.5 | 三次连测 |

已知边界：200K decode 30 t/s 在当前模型+草稿+硬件下暂无可证路径（attention 已贴字节地板，
KV 读 13.9 GB/step）；保障读数 = 200K 21.9–22.5 t/s。

## 平台要求

- NVIDIA DRIVE Thor 类域控（sm_101a），DriveOS 7.0.3，CUDA 12.8 用户态
- GPU 大页池 ≥46 GiB（23552 页 × 2MiB，冷启动分配；宿主机准备脚本见 GitHub 仓库）
- 板上运行时路径含 tmpfs，重启后二进制需重传或放持久分区

## 许可

Apache-2.0（本包脚本与文档）；上游 llama.cpp 遵循其自身许可；模型权重各归其上游许可。
