---
title: "PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity"
canonical_url: "https://www.modelscope.cn/papers/125501"
md_url: "https://www.modelscope.cn/papers/125501.md"
arxiv_id: 2503.07677
published: 2025-03-10
last_updated: 2025-03-10
authors:
  - "Kwanyoung Kim"
  - "Byeongsu Sim"
model_name: PLADIS
model_developer: "三星研究院"
domain:
  - "计算机视觉"
  - "深度学习"
type:
  - "计算机视觉"
  - "深度学习"
  - "Machine Learning (cs.LG)"
  - "Artificial Intelligence (cs.AI)"
arxiv_url: "https://arxiv.org/abs/2503.07677"
pdf_url: "https://arxiv.org/pdf/2503.07677.pdf"
---

# PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity

> Diffusion models have shown impressive results in generating high-quality conditional samples using guidance techniques such as Classifier-Free Guidance (CFG). However, existing methods often require additional training or neural function evaluations (NFEs),…

「PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity」是 ModelScope 魔搭社区收录的论文，arXiv 2503.07677，作者为 Kwanyoung Kim, Byeongsu Sim，发表于 2025-03-10，属于 计算机视觉、深度学习 领域。

- **ArXiv**: 2503.07677
- **Published**: 2025-03-10
- **Authors**: Kwanyoung Kim, Byeongsu Sim
- **Model**: PLADIS
- **Developer**: 三星研究院
- **Domain**: 计算机视觉, 深度学习
- **ArXiv URL**: https://arxiv.org/abs/2503.07677
- **PDF**: https://arxiv.org/pdf/2503.07677.pdf

Source: https://www.modelscope.cn/papers/125501

---

> PLADIS：利用稀疏注意力机制突破扩散模型推理极限

## 摘要

本文探讨了在扩散模型中利用稀疏注意力机制提升推理阶段性能的方法。现有的引导技术如Classifier-Free Guidance（CFG）虽然有效，但需要额外的训练或神经函数评估（NFEs），并且不适用于引导蒸馏模型。为了解决这些问题，作者提出了名为PLADIS的新方法，该方法通过在推理过程中使用稀疏注意力机制来增强预训练模型（如U-Net和Transformer）。具体来说，PLADIS利用softmax及其稀疏版本在交叉注意力层中推断query-key相关性，无需额外训练或NFEs。实验表明，PLADIS显著改善了文本对齐效果和生成图像的质量，并且可以无缝集成到各种引导技术和引导蒸馏模型中。此外，PLADIS还展示了现代Hopfield网络和稀疏Hopfield网络的噪声鲁棒性优势。

## Abstract

Diffusion models have shown impressive results in generating high-quality conditional samples using guidance techniques such as Classifier-Free Guidance (CFG). However, existing methods often require additional training or neural function evaluations (NFEs), making them incompatible with guidance-distilled models. Also, they rely on heuristic approaches that need identifying target layers. In this work, we propose a novel and efficient method, termed PLADIS, which boosts pre-trained models (U-Net/Transformer) by leveraging sparse attention. Specifically, we extrapolate query-key correlations using softmax and its sparse counterpart in the cross-attention layer during inference, without requiring extra training or NFEs. By leveraging the noise robustness of sparse attention, our PLADIS unleashes the latent potential of text-to-image diffusion models, enabling them to excel in areas where they once struggled with newfound effectiveness. It integrates seamlessly with guidance techniques, including guidance-distilled models. Extensive experiments show notable improvements in text alignment and human preference, offering a highly efficient and universally applicable solution.
