---
title: "TPDiff: Temporal Pyramid Video Diffusion Model"
canonical_url: "https://www.modelscope.cn/papers/125911"
md_url: "https://www.modelscope.cn/papers/125911.md"
arxiv_id: 2503.09566
published: 2025-03-12
last_updated: 2025-03-12
authors:
  - "Lingmin Ran"
  - "Mike Zheng Shou"
model_name: TPDiff
model_developer: "新加坡国立大学显示实验室"
domain:
  - "计算机视觉"
  - "深度学习"
type:
  - "计算机视觉"
  - "深度学习"
  - "Computer Vision and Pattern Recognition (cs.CV)"
arxiv_url: "https://arxiv.org/abs/2503.09566"
pdf_url: "https://arxiv.org/pdf/2503.09566.pdf"
---

# TPDiff: Temporal Pyramid Video Diffusion Model

> The development of video diffusion models unveils a significant challenge: the substantial computational demands. To mitigate this challenge, we note that the reverse process of diffusion exhibits an inherent entropy-reducing nature. Given the inter-frame…

「TPDiff: Temporal Pyramid Video Diffusion Model」是 ModelScope 魔搭社区收录的论文，arXiv 2503.09566，作者为 Lingmin Ran, Mike Zheng Shou，发表于 2025-03-12，属于 计算机视觉、深度学习 领域。

- **ArXiv**: 2503.09566
- **Published**: 2025-03-12
- **Authors**: Lingmin Ran, Mike Zheng Shou
- **Model**: TPDiff
- **Developer**: 新加坡国立大学显示实验室
- **Domain**: 计算机视觉, 深度学习
- **ArXiv URL**: https://arxiv.org/abs/2503.09566
- **PDF**: https://arxiv.org/pdf/2503.09566.pdf

Source: https://www.modelscope.cn/papers/125911

---

> TPDiff：用时间金字塔优化视频扩散模型的高效利器

## 摘要

本文针对视频扩散模型的高计算需求问题，提出了一种名为TPDiff（Temporal Pyramid Video Diffusion Model）的方法。研究背景显示，尽管视频扩散模型在艺术创作、机器人技术及虚拟现实等领域表现出巨大潜力，但其训练和推理成本因需要同时建模时空分布而异常高昂。TPDiff通过引入时间金字塔结构，在扩散过程中逐步增加帧率，仅在最后阶段使用全帧率，从而显著优化了计算效率。此外，为训练多阶段扩散模型，作者设计了分阶段扩散（stage-wise diffusion）框架，解决了不同帧率阶段间的连续性问题，并适用于多种扩散形式。实验结果表明，该方法相比传统扩散模型降低了50%的训练成本，并将推理效率提升了1.5倍。TPDiff不仅有效缓解了视频生成中的计算瓶颈，还展示了对不同扩散形式的通用性。

## Abstract

The development of video diffusion models unveils a significant challenge: the substantial computational demands. To mitigate this challenge, we note that the reverse process of diffusion exhibits an inherent entropy-reducing nature. Given the inter-frame redundancy in video modality, maintaining full frame rates in high-entropy stages is unnecessary. Based on this insight, we propose TPDiff, a unified framework to enhance training and inference efficiency. By dividing diffusion into several stages, our framework progressively increases frame rate along the diffusion process with only the last stage operating on full frame rate, thereby optimizing computational efficiency. To train the multi-stage diffusion model, we introduce a dedicated training framework: stage-wise diffusion. By solving the partitioned probability flow ordinary differential equations (ODE) of diffusion under aligned data and noise, our training strategy is applicable to various diffusion forms and further enhances training efficiency. Comprehensive experimental evaluations validate the generality of our method, demonstrating 50% reduction in training cost and 1.5x improvement in inference efficiency.
