---
title: "VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control"
canonical_url: "https://www.modelscope.cn/papers/124052"
md_url: "https://www.modelscope.cn/papers/124052.md"
arxiv_id: 2503.05639
published: 2025-03-07
last_updated: 2025-03-07
authors:
  - "Yuxuan Bian"
  - "Zhaoyang Zhang"
  - "Xuan Ju"
  - "Mingdeng Cao"
  - "Liangbin Xie"
  - "Ying Shan"
  - "Qiang Xu"
model_name: VideoPainter
model_developer: "香港中文大学, 腾讯ARC实验室, 澳门大学, 东京大学"
domain:
  - "计算机视觉"
  - "深度学习"
type:
  - "计算机视觉"
  - "深度学习"
  - "Computer Vision and Pattern Recognition (cs.CV)"
  - "Artificial Intelligence (cs.AI)"
  - "Multimedia (cs.MM)"
arxiv_url: "https://arxiv.org/abs/2503.05639"
pdf_url: "https://arxiv.org/pdf/2503.05639.pdf"
---

# VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control

> Video inpainting, which aims to restore corrupted video content, has experienced substantial progress. Despite these advances, existing methods, whether propagating unmasked region pixels through optical flow and receptive field priors, or extending…

「VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control」是 ModelScope 魔搭社区收录的论文，arXiv 2503.05639，作者为 Yuxuan Bian, Zhaoyang Zhang, Xuan Ju et al.，发表于 2025-03-07，属于 计算机视觉、深度学习 领域。

- **ArXiv**: 2503.05639
- **Published**: 2025-03-07
- **Authors**: Yuxuan Bian, Zhaoyang Zhang, Xuan Ju, Mingdeng Cao, Liangbin Xie, Ying Shan, Qiang Xu
- **Model**: VideoPainter
- **Developer**: 香港中文大学, 腾讯ARC实验室, 澳门大学, 东京大学
- **Domain**: 计算机视觉, 深度学习
- **ArXiv URL**: https://arxiv.org/abs/2503.05639
- **PDF**: https://arxiv.org/pdf/2503.05639.pdf

Source: https://www.modelscope.cn/papers/124052

---

> VideoPainter：任意长度视频修复与编辑的即插即用上下文控制

## 摘要

本文提出了一种名为VideoPainter的双分支框架，用于任意长度视频的修复和编辑。现有方法在处理完全遮挡对象、平衡背景保留与前景生成以及保持长时间视频中物体身份一致性方面存在挑战。VideoPainter通过引入轻量级上下文编码器，将背景引导注入任何预训练的视频扩散变换器中，实现了高效且密集的背景控制和用户自定义控制。该框架还引入了新的区域重采样技术，以确保任意长度视频修复中的ID一致性。此外，作者开发了一个可扩展的数据集管道，构建了VPData和VPBench——目前最大的视频修复数据集，包含超过390K个带有分割掩码和密集字幕的片段。实验结果表明，VideoPainter在视频质量、掩码区域保留和文本对齐等8个关键指标上均达到了最先进的性能。

## Abstract

Video inpainting, which aims to restore corrupted video content, has experienced substantial progress. Despite these advances, existing methods, whether propagating unmasked region pixels through optical flow and receptive field priors, or extending image-inpainting models temporally, face challenges in generating fully masked objects or balancing the competing objectives of background context preservation and foreground generation in one model, respectively. To address these limitations, we propose a novel dual-stream paradigm VideoPainter that incorporates an efficient context encoder (comprising only 6% of the backbone parameters) to process masked videos and inject backbone-aware background contextual cues to any pre-trained video DiT, producing semantically consistent content in a plug-and-play manner. This architectural separation significantly reduces the model's learning complexity while enabling nuanced integration of crucial background context. We also introduce a novel target region ID resampling technique that enables any-length video inpainting, greatly enhancing our practical applicability. Additionally, we establish a scalable dataset pipeline leveraging current vision understanding models, contributing VPData and VPBench to facilitate segmentation-based inpainting training and assessment, the largest video inpainting dataset and benchmark to date with over 390K diverse clips. Using inpainting as a pipeline basis, we also explore downstream applications including video editing and video editing pair data generation, demonstrating competitive performance and significant practical potential. Extensive experiments demonstrate VideoPainter's superior performance in both any-length video inpainting and editing, across eight key metrics, including video quality, mask region preservation, and textual coherence.
