---
title: flash-attn-windows-blackwell
canonical_url: "https://www.modelscope.cn/models/wxd11011/flash-attn-windows-blackwell"
md_url: "https://www.modelscope.cn/models/wxd11011/flash-attn-windows-blackwell.md"
repository: wxd11011/flash-attn-windows-blackwell
chinese_name: "Flash-Attention 预编译包｜Windows + RTX50系(Blackwell) + CUDA13.0"
last_updated: 2026-08-03
library_name:
  - pytorch
downloads: 0
stars: 0
---

# flash-attn-windows-blackwell

> flash-attn-windows-blackwell - wxd11011 在 ModelScope 开源的模型。flash-attn 2.8.4/2.8.3 Windows预编译Wheel，支持RTX50系(Blackwell SM120)+CUDA13.0，覆盖PyTorch 2.9/2.13+Python 3.12/3.13，即下即用。

- **Repository**: wxd11011/flash-attn-windows-blackwell
- **Downloads**: 0
- **Stars**: 0
- **Last updated**: 2026-08-03

Source: https://www.modelscope.cn/models/wxd11011/flash-attn-windows-blackwell

---

# Flash-Attention 预编译 Wheel 合集（Windows + Blackwell 专用）

## 🎯 简介

本项目提供 **flash-attn** 在 Windows 平台下的预编译 wheel，专门针对 **NVIDIA RTX 50 系列（Blackwell 架构 / SM 12.0）** 显卡编译，基于 **CUDA 13.0**。

特别适合 **ComfyUI / Stable Diffusion / Flux / LLM 推理** 等场景的用户**直接下载使用，无需自行配置编译环境**。

本仓库提供 **两个版本**，请根据你已安装的 PyTorch 和 Python 版本选择：

| 版本 | flash-attn | Python | PyTorch | 文件名 | SHA256 |
|------|-----------|--------|---------|--------|---------|
| 🅰️ | 2.8.4 | 3.13 | 2.13.0+cu130 | `flash_attn-2.8.4+...cp313...whl` |5ed057fd53fbd870fedb7f309b0f1624196e47ce237311a05c46b112bf089744|
| 🅱️ | 2.8.3 | 3.12 | 2.9.1+cu130 | `flash_attn-2.8.3+...cp312...whl` |c3e10b874565b64e18919c90e2ab9af5f7bc9d97296fc052991c2686533487a5|

---

## ⚠️ 如何选择版本（重要）

**wheel 对版本极其敏感，必须与你的环境完全一致。** 选择依据是你**当前已安装的 PyTorch 版本**。

1. 先检查你的环境：
   ```bash
   python -c "import torch, sys; print('Python:', sys.version); print('PyTorch:', torch.__version__, '| CUDA:', torch.version.cuda)"
   ```

2. 对照选择：
   - 输出 `Python 3.13` + `PyTorch 2.13.0+cu130` → 下载 **版本 🅰️（flash-attn 2.8.4）**
   - 输出 `Python 3.12` + `PyTorch 2.9.1+cu130` → 下载 **版本 🅱️（flash-attn 2.8.3）**

> ❗ 如果版本不匹配，安装时会报 `not supported wheel on this platform`，或导入时报 `undefined symbol` / `DLL load failed`。

---

## 📦 安装方法

下载对应的 `.whl` 文件后，在对应 Python 环境中执行：

**版本 🅰️（Python 3.13 + PyTorch 2.13）：**
```bash
pip install flash_attn-2.8.4+cu13torch2.13cxx11abiTRUE-cp313-cp313-win_amd64.whl
```

**版本 🅱️（Python 3.12 + PyTorch 2.9.1）：**
```bash
pip install flash_attn-2.8.3+cu13torch2.9.1cxx11abiTRUE-cp312-cp312-win_amd64.whl
```

**ComfyUI 便携版用户**，请使用内置 Python 的完整路径：
```bash
ComfyUI_windows_portable_nvidia\python_embeded\python.exe -m pip install 对应的whl文件
```

---

## ✅ 安装验证

```python
import torch
from flash_attn import flash_attn_func
import flash_attn

def test_flash_attn():
    print(f"flash_attn version: {flash_attn.__version__}")

    if not torch.cuda.is_available():
        print("❌ CUDA 不可用，Flash Attention 需要 GPU 支持。")
        return False

    # 设置小尺寸参数，避免显存溢出
    batch_size, seq_len, n_heads, head_dim = 2, 8, 4, 32
    dtype = torch.float16
    device = "cuda"

    # 随机生成 q, k, v（带梯度）
    q = torch.randn((batch_size, seq_len, n_heads, head_dim), dtype=dtype, device=device, requires_grad=True)
    k = torch.randn((batch_size, seq_len, n_heads, head_dim), dtype=dtype, device=device, requires_grad=True)
    v = torch.randn((batch_size, seq_len, n_heads, head_dim), dtype=dtype, device=device, requires_grad=True)

    # 前向传播
    out = flash_attn_func(q, k, v, causal=False)

    # 检查输出形状
    expected = (batch_size, seq_len, n_heads, head_dim)
    if out.shape != expected:
        print(f"❌ 输出形状错误: {out.shape} != {expected}")
        return False

    # 反向传播（测试梯度）
    loss = out.sum()
    loss.backward()

    if q.grad is None or k.grad is None or v.grad is None:
        print("❌ 梯度计算失败")
        return False

    print("✅ Flash Attention 测试通过！")
    return True

if __name__ == "__main__":
    test_flash_attn()
```

---

## 🎮 支持的 GPU 架构

### 版本 🅰️（flash-attn 2.8.4）
支持 **5 种架构**：

| 架构 | 对应 GPU |
|------|---------|
| sm_80 | A100、RTX 3090 |
| sm_90 | H100、H200 |
| sm_100 | B100、B200（数据中心） |
| sm_110 | Blackwell 变体 |
| sm_120 | **RTX 5060/5070/5080/5090** ✅ |

> 📌 版本 🅰️ 的 sm_120 包含 PTX 中间码，未来 SM 12.x 新 GPU 也可通过 JIT 运行，具备前向兼容能力。

### 版本 🅱️（flash-attn 2.8.3）
支持 **4 种架构**：

| 架构 | 对应 GPU |
|------|---------|
| sm_80 | A100、RTX 3090 |
| sm_90 | H100、H200 |
| sm_100 | B100、B200（数据中心） |
| sm_120 | **RTX 5060/5070/5080/5090** ✅ |

### ❌ 两个版本均不支持的 GPU
- **sm_86**：RTX 3060 / 3070 / 3080 / 3080 Ti / A10 / A40
- **sm_89**：RTX 4060 / 4070 / 4080 / 4090 / L40 / L40S

> 如果你的 GPU 属于上述"不支持"列表，本 wheel **无法使用**，请勿下载。

---

## 📌 适用场景

- ✅ ComfyUI 自定义节点加速（如需要 flash-attn 的节点）
- ✅ Stable Diffusion / Flux 等大模型推理加速
- ✅ LLM 推理（vLLM / Transformers 等依赖 flash-attn 的框架）
- ✅ 任何在 Windows + Blackwell GPU 上需要 Flash Attention 的项目

---

## ⚠️ 注意事项

1. **版本必须严格匹配**，尤其是 PyTorch 版本和 Python 版本
2. 如果你的 GPU 不是 Blackwell 架构（SM 12.x），请确认 wheel 中包含你 GPU 的架构代码
3. 本 wheel 为自行编译，**非官方发布**，请从可信渠道下载并自行校验
4. 若导入时报 `DLL load failed`，请确认已安装 **CUDA 13.0 运行时** 和 **VC++ 2022 运行库**

---

## 🙏 致谢与来源

- Flash-Attention 官方仓库：https://github.com/Dao-AILab/flash-attention
- PyTorch：https://pytorch.org
- CUDA Toolkit：https://developer.nvidia.com/cuda-toolkit
- 编译参考文章：AITechLab ——[Windows 下成功编译 Flash Attention 2.8.3 （flash-attn /flash_attn）个人复盘记录](https://blog.csdn.net/u014451778/article/details/156356011)（CSDN）

> 本项目编译过程中参考了上述 CSDN 文章的思路与方法，在此向原作者 **AITechLab** 表示感谢。

---

## 📄 许可

本项目编译产物遵循 flash-attn 原项目的 **BSD-3-Clause** 许可证。

如有问题欢迎在评论区留言～
