---
title: "Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models"
canonical_url: "https://www.modelscope.cn/papers/126168"
md_url: "https://www.modelscope.cn/papers/126168.md"
arxiv_id: 2503.09573
published: 2025-03-12
last_updated: 2025-03-12
authors:
  - "Marianne Arriola"
  - "Aaron Gokaslan"
  - "Justin T Chiu"
  - "Zhihan Yang"
  - "Zhixuan Qi"
  - "Jiaqi Han"
  - "Subham Sekhar Sahoo"
  - "Volodymyr Kuleshov"
model_name: "BD3-LMs (Block Discrete Denoising Diffusion Language Models)"
model_developer: "康奈尔科技学院，斯坦福大学，Cohere"
domain:
  - "自然语言处理"
  - "深度学习"
  - "机器学习"
type:
  - "自然语言处理"
  - "深度学习"
  - "机器学习"
  - "Machine Learning (cs.LG)"
  - "Artificial Intelligence (cs.AI)"
arxiv_url: "https://arxiv.org/abs/2503.09573"
pdf_url: "https://arxiv.org/pdf/2503.09573.pdf"
---

# Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models

> Diffusion language models offer unique benefits over autoregressive models due to their potential for parallelized generation and controllability, yet they lag in likelihood modeling and are limited to fixed-length generation. In this work, we introduce a…

「Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models」是 ModelScope 魔搭社区收录的论文，arXiv 2503.09573，作者为 Marianne Arriola, Aaron Gokaslan, Justin T Chiu et al.，发表于 2025-03-12，属于 自然语言处理、深度学习、机器学习 领域。

- **ArXiv**: 2503.09573
- **Published**: 2025-03-12
- **Authors**: Marianne Arriola, Aaron Gokaslan, Justin T Chiu, Zhihan Yang, Zhixuan Qi, Jiaqi Han, Subham Sekhar Sahoo, Volodymyr Kuleshov
- **Model**: BD3-LMs (Block Discrete Denoising Diffusion Language Models)
- **Developer**: 康奈尔科技学院，斯坦福大学，Cohere
- **Domain**: 自然语言处理, 深度学习, 机器学习
- **ArXiv URL**: https://arxiv.org/abs/2503.09573
- **PDF**: https://arxiv.org/pdf/2503.09573.pdf

Source: https://www.modelscope.cn/papers/126168

---

> 块扩散语言模型：融合自回归与扩散优势的全新生成范式

## 摘要

本文旨在解决离散扩散模型在生成任意长度序列、推理效率和质量上的局限性。传统的自回归模型虽然生成高质量的文本，但无法并行化生成且固定长度；而扩散模型虽然支持并行生成，但质量和灵活性较低。为克服这些缺点，本文提出了一种新的模型——块离散去噪扩散语言模型（BD3-LMs）。该模型结合了自回归模型和扩散模型的优点，通过块级自回归分布和块内扩散过程实现灵活长度生成，并利用KV缓存和并行采样提高推理效率。

具体而言，BD3-LMs将序列划分为多个块，每个块内的条件概率由离散扩散模型定义，块间则采用自回归方式建模。作者提出了高效的训练算法，包括梯度方差估计器和数据驱动的噪声调度方法，以降低扩散目标的梯度方差，从而缩小与自回归模型之间的困惑度差距。实验结果表明，BD3-LMs在语言建模基准测试中取得了离散扩散模型的新状态最优困惑度，并能够生成超过训练上下文长度的任意长度序列。此外，与基于高斯扩散的半自回归模型相比，BD3-LMs具有可追踪的似然估计和更少的生成步骤。

总体而言，本文通过引入块扩散模型及其优化训练方法，显著提升了离散扩散模型的性能，为其在自然语言生成任务中的应用提供了新方向。

## Abstract

Diffusion language models offer unique benefits over autoregressive models due to their potential for parallelized generation and controllability, yet they lag in likelihood modeling and are limited to fixed-length generation. In this work, we introduce a class of block diffusion language models that interpolate between discrete denoising diffusion and autoregressive models. Block diffusion overcomes key limitations of both approaches by supporting flexible-length generation and improving inference efficiency with KV caching and parallel token sampling. We propose a recipe for building effective block diffusion models that includes an efficient training algorithm, estimators of gradient variance, and data-driven noise schedules to minimize the variance. Block diffusion sets a new state-of-the-art performance among diffusion models on language modeling benchmarks and enables generation of arbitrary-length sequences. We provide the code, along with the model weights and blog post on the project page: this https URL
