---
title: "Forgetting Transformer: Softmax Attention with a Forget Gate"
canonical_url: "https://www.modelscope.cn/papers/122938"
md_url: "https://www.modelscope.cn/papers/122938.md"
arxiv_id: 2503.02130
published: 2025-03-03
last_updated: 2025-03-03
authors:
  - "Zhixuan Lin"
  - "Evgenii Nikishin"
  - "Xu Owen He"
  - "Aaron Courville"
model_name: "Forgetting Transformer (FoX)"
model_developer: "Mila与蒙特利尔大学"
domain:
  - "自然语言处理"
  - "深度学习"
type:
  - "自然语言处理"
  - "深度学习"
  - "Machine Learning (cs.LG)"
  - "Artificial Intelligence (cs.AI)"
  - "Computation and Language (cs.CL)"
arxiv_url: "https://arxiv.org/abs/2503.02130"
pdf_url: "https://arxiv.org/pdf/2503.02130.pdf"
code_link: "https://github.com/zhixuan-lin/forgetting-transformer"
---

# Forgetting Transformer: Softmax Attention with a Forget Gate

> An essential component of modern recurrent sequence models is the forget gate. While Transformers do not have an explicit recurrent form, we show that a forget gate can be naturally incorporated into Transformers by down-weighting the unnormalized attention…

「Forgetting Transformer: Softmax Attention with a Forget Gate」是 ModelScope 魔搭社区收录的论文，arXiv 2503.02130，作者为 Zhixuan Lin, Evgenii Nikishin, Xu Owen He et al.，发表于 2025-03-03，属于 自然语言处理、深度学习 领域。

- **ArXiv**: 2503.02130
- **Published**: 2025-03-03
- **Authors**: Zhixuan Lin, Evgenii Nikishin, Xu Owen He, Aaron Courville
- **Model**: Forgetting Transformer (FoX)
- **Developer**: Mila与蒙特利尔大学
- **Domain**: 自然语言处理, 深度学习
- **ArXiv URL**: https://arxiv.org/abs/2503.02130
- **PDF**: https://arxiv.org/pdf/2503.02130.pdf
- **Code**: https://github.com/zhixuan-lin/forgetting-transformer

Source: https://www.modelscope.cn/papers/122938

---

> 引入遗忘门机制：Forgetting Transformer在长上下文建模中的突破

## 摘要

本文提出了一种新的Transformer变体——Forgetting Transformer（简称FoX），该模型通过引入遗忘门机制来增强Transformer在长上下文建模中的表现。传统的Transformer缺乏显式的遗忘机制，而这种机制在递归序列模型中已被证明对短上下文任务至关重要。作者通过调整未归一化的注意力分数，将遗忘门机制引入到Transformer中，称为Forgetting Attention。实验结果表明，FoX在长上下文语言建模、长度外推和短上下文下游任务上均优于标准Transformer，并且在长上下文任务上的表现与Transformer相当。此外，FoX无需使用位置嵌入，并且可以通过简单的修改与FlashAttention算法兼容。为了进一步提升性能，作者还引入了一种改进的“Pro”块设计，结合了递归序列模型中常用的架构组件。通过对不同验证上下文长度的困惑度和每词损失的分析，验证了FoX在处理长上下文信息方面的优越性。

## Abstract

An essential component of modern recurrent sequence models is the forget gate. While Transformers do not have an explicit recurrent form, we show that a forget gate can be naturally incorporated into Transformers by down-weighting the unnormalized attention scores in a data-dependent way. We name this attention mechanism the Forgetting Attention and the resulting model the Forgetting Transformer (FoX). We show that FoX outperforms the Transformer on long-context language modeling, length extrapolation, and short-context downstream tasks, while performing on par with the Transformer on long-context downstream tasks. Moreover, it is compatible with the FlashAttention algorithm and does not require any positional embeddings. Several analyses, including the needle-in-the-haystack test, show that FoX also retains the Transformer's superior long-context capabilities over recurrent sequence models such as Mamba-2, HGRN2, and DeltaNet. We also introduce a "Pro" block design that incorporates some common architectural components in recurrent sequence models and find it significantly improves the performance of both FoX and the Transformer. Our code is available at this https URL.
