---
title: "ReCamMaster: Camera-Controlled Generative Rendering from A Single Video"
canonical_url: "https://www.modelscope.cn/papers/127016"
md_url: "https://www.modelscope.cn/papers/127016.md"
arxiv_id: 2503.11647
published: 2025-03-14
last_updated: 2025-03-14
authors:
  - "Jianhong Bai"
  - "Menghan Xia"
  - "Xiao Fu"
  - "Xintao Wang"
  - "Lianrui Mu"
  - "Jinwen Cao"
  - "Zuozhu Liu"
  - "Haoji Hu"
  - "Xiang Bai"
  - "Pengfei Wan"
  - "Di Zhang"
model_name: ReCamMaster
model_developer: "浙江大学，快手科技，香港中文大学，华中科技大学"
domain:
  - "计算机视觉"
  - "深度学习"
type:
  - "计算机视觉"
  - "深度学习"
  - "Computer Vision and Pattern Recognition (cs.CV)"
arxiv_url: "https://arxiv.org/abs/2503.11647"
pdf_url: "https://arxiv.org/pdf/2503.11647.pdf"
code_link: "https://jianhongbai.github.io/ReCamMaster/"
---

# ReCamMaster: Camera-Controlled Generative Rendering from A Single Video

> Camera control has been actively studied in text or image conditioned video generation tasks. However, altering camera trajectories of a given video remains under-explored, despite its importance in the field of video creation. It is non-trivial due to the…

「ReCamMaster: Camera-Controlled Generative Rendering from A Single Video」是 ModelScope 魔搭社区收录的论文，arXiv 2503.11647，作者为 Jianhong Bai, Menghan Xia, Xiao Fu et al.，发表于 2025-03-14，属于 计算机视觉、深度学习 领域。

- **ArXiv**: 2503.11647
- **Published**: 2025-03-14
- **Authors**: Jianhong Bai, Menghan Xia, Xiao Fu, Xintao Wang, Lianrui Mu, Jinwen Cao, Zuozhu Liu, Haoji Hu, Xiang Bai, Pengfei Wan, Di Zhang
- **Model**: ReCamMaster
- **Developer**: 浙江大学，快手科技，香港中文大学，华中科技大学
- **Domain**: 计算机视觉, 深度学习
- **ArXiv URL**: https://arxiv.org/abs/2503.11647
- **PDF**: https://arxiv.org/pdf/2503.11647.pdf
- **Code**: https://jianhongbai.github.io/ReCamMaster/

Source: https://www.modelscope.cn/papers/127016

---

> ReCamMaster：重塑视频视角的艺术大师

## 摘要

本研究针对视频生成任务中相机轨迹控制的挑战，提出了一种名为ReCamMaster的新框架。研究背景方面，相机运动在影视制作中至关重要，但业余摄像师因硬件和技术限制难以实现专业级效果，因此需要一种后处理方法来优化相机轨迹。ReCamMaster通过结合预训练的文本到视频生成模型和一种新颖的视频条件机制，实现了对任意视频在新相机轨迹下的动态场景再生。为了解决合格训练数据稀缺的问题，作者使用Unreal Engine 5构建了一个高质量的多相机同步视频数据集，包含13.6K动态场景、40个高质3D环境和122K不同相机轨迹。此外，作者还设计了细致的训练策略以提高模型对多样输入的鲁棒性。

提出的方法上，ReCamMaster的核心创新在于利用预训练文本到视频扩散模型的生成能力，并通过精心设计的双条件框架（源视频和目标相机轨迹）进行约束。为了实现这一目标，作者提出了一个条件视频注入机制，包括帧维度、通道维度和视图维度的条件化方法。实验结果表明，ReCamMaster显著优于现有方法和强基线，在视频稳定、超分辨率和外画等实际应用中表现出巨大潜力。

取得的结果和价值贡献方面，该研究不仅提供了高质量的多相机同步视频数据集以推动相关领域的研究，还深入验证了一种有效的视频条件机制。广泛的实验表明，ReCamMaster大幅提升了视频重生成技术的水平，并展示了其在多个现实场景中的应用前景。

## Abstract

Camera control has been actively studied in text or image conditioned video generation tasks. However, altering camera trajectories of a given video remains under-explored, despite its importance in the field of video creation. It is non-trivial due to the extra constraints of maintaining multiple-frame appearance and dynamic synchronization. To address this, we present ReCamMaster, a camera-controlled generative video re-rendering framework that reproduces the dynamic scene of an input video at novel camera trajectories. The core innovation lies in harnessing the generative capabilities of pre-trained text-to-video models through a simple yet powerful video conditioning mechanism -- its capability often overlooked in current research. To overcome the scarcity of qualified training data, we construct a comprehensive multi-camera synchronized video dataset using Unreal Engine 5, which is carefully curated to follow real-world filming characteristics, covering diverse scenes and camera movements. It helps the model generalize to in-the-wild videos. Lastly, we further improve the robustness to diverse inputs through a meticulously designed training strategy. Extensive experiments tell that our method substantially outperforms existing state-of-the-art approaches and strong baselines. Our method also finds promising applications in video stabilization, super-resolution, and outpainting. Project page: this https URL
