---
title: "Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers"
canonical_url: "https://www.modelscope.cn/papers/126998"
md_url: "https://www.modelscope.cn/papers/126998.md"
arxiv_id: 2503.11579
published: 2025-03-14
last_updated: 2025-03-14
authors:
  - "Weiming Ren"
  - "Wentao Ma"
  - "Huan Yang"
  - "Cong Wei"
  - "Ge Zhang"
  - "Wenhu Chen"
model_name: VAMBA
model_developer: "滑铁卢大学，多伦多大学，01.AI，向量研究所，M-A-P"
domain:
  - "自然语言处理"
  - "计算机视觉"
  - "深度学习"
type:
  - "自然语言处理"
  - "计算机视觉"
  - "深度学习"
  - "Computer Vision and Pattern Recognition (cs.CV)"
arxiv_url: "https://arxiv.org/abs/2503.11579"
pdf_url: "https://arxiv.org/pdf/2503.11579.pdf"
---

# Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers

> State-of-the-art transformer-based large multimodal models (LMMs) struggle to handle hour-long video inputs due to the quadratic complexity of the causal self-attention operations, leading to high computational costs during training and inference. Existing…

「Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers」是 ModelScope 魔搭社区收录的论文，arXiv 2503.11579，作者为 Weiming Ren, Wentao Ma, Huan Yang et al.，发表于 2025-03-14，属于 自然语言处理、计算机视觉、深度学习 领域。

- **ArXiv**: 2503.11579
- **Published**: 2025-03-14
- **Authors**: Weiming Ren, Wentao Ma, Huan Yang, Cong Wei, Ge Zhang, Wenhu Chen
- **Model**: VAMBA
- **Developer**: 滑铁卢大学，多伦多大学，01.AI，向量研究所，M-A-P
- **Domain**: 自然语言处理, 计算机视觉, 深度学习
- **ArXiv URL**: https://arxiv.org/abs/2503.11579
- **PDF**: https://arxiv.org/pdf/2503.11579.pdf

Source: https://www.modelscope.cn/papers/126998

---

> VAMBA：用混合Mamba-Transformer革新长时间视频理解

## 摘要

本文研究背景聚焦于当前基于Transformer的大型多模态模型（LMMs）在处理长时间视频输入时面临的挑战。由于因果自注意力操作的二次复杂度，这些模型在训练和推理过程中计算成本极高。现有基于令牌压缩的方法虽减少视频令牌数量，但可能导致信息丢失且效率不足。为此，本文提出了一种名为VAMBA的混合Mamba-Transformer模型，用于高效处理长达一小时的视频理解任务。VAMBA通过使用Mamba-2块以线性复杂度编码视频令牌，无需任何令牌减少即可在单个GPU上处理超过1024帧（640×360）的视频数据。实验结果表明，在训练和推理过程中，VAMBA至少减少了50%的GPU内存使用，并将每步训练速度几乎提高一倍。与先前高效的视频LMMs相比，VAMBA在具有挑战性的长时间视频理解基准LVBench上准确率提高了4.3%，并在其他长短视频理解任务中保持了强劲性能。此外，作者还进行了全面的消融研究，验证了采用Mamba-2块和从预训练自注意力层初始化交叉注意力权重对高性能实现的关键作用。

## Abstract

State-of-the-art transformer-based large multimodal models (LMMs) struggle to handle hour-long video inputs due to the quadratic complexity of the causal self-attention operations, leading to high computational costs during training and inference. Existing token compression-based methods reduce the number of video tokens but often incur information loss and remain inefficient for extremely long sequences. In this paper, we explore an orthogonal direction to build a hybrid Mamba-Transformer model (VAMBA) that employs Mamba-2 blocks to encode video tokens with linear complexity. Without any token reduction, VAMBA can encode more than 1024 frames (640$\times$360) on a single GPU, while transformer-based models can only encode 256 frames. On long video input, VAMBA achieves at least 50% reduction in GPU memory usage during training and inference, and nearly doubles the speed per training step compared to transformer-based LMMs. Our experimental results demonstrate that VAMBA improves accuracy by 4.3% on the challenging hour-long video understanding benchmark LVBench over prior efficient video LMMs, and maintains strong performance on a broad spectrum of long and short video understanding tasks.
