---
title: "Reangle-A-Video: 4D Video Generation as Video-to-Video Translation"
canonical_url: "https://www.modelscope.cn/papers/126077"
md_url: "https://www.modelscope.cn/papers/126077.md"
arxiv_id: 2503.09151
published: 2025-03-12
last_updated: 2025-03-12
authors:
  - "Hyeonho Jeong"
  - "Suhyeon Lee"
  - "Jong Chul Ye"
model_name: Reangle-A-Video
model_developer: "KAIST人工智能研究院"
domain:
  - "计算机视觉"
  - "深度学习"
type:
  - "计算机视觉"
  - "深度学习"
  - "Computer Vision and Pattern Recognition (cs.CV)"
  - "Artificial Intelligence (cs.AI)"
arxiv_url: "https://arxiv.org/abs/2503.09151"
pdf_url: "https://arxiv.org/pdf/2503.09151.pdf"
code_link: "https://hyeonho99.github.io/reangle-a-video/"
---

# Reangle-A-Video: 4D Video Generation as Video-to-Video Translation

> We introduce Reangle-A-Video, a unified framework for generating synchronized multi-view videos from a single input video. Unlike mainstream approaches that train multi-view video diffusion models on large-scale 4D datasets, our method reframes the…

「Reangle-A-Video: 4D Video Generation as Video-to-Video Translation」是 ModelScope 魔搭社区收录的论文，arXiv 2503.09151，作者为 Hyeonho Jeong, Suhyeon Lee, Jong Chul Ye，发表于 2025-03-12，属于 计算机视觉、深度学习 领域。

- **ArXiv**: 2503.09151
- **Published**: 2025-03-12
- **Authors**: Hyeonho Jeong, Suhyeon Lee, Jong Chul Ye
- **Model**: Reangle-A-Video
- **Developer**: KAIST人工智能研究院
- **Domain**: 计算机视觉, 深度学习
- **ArXiv URL**: https://arxiv.org/abs/2503.09151
- **PDF**: https://arxiv.org/pdf/2503.09151.pdf
- **Code**: https://hyeonho99.github.io/reangle-a-video/

Source: https://www.modelscope.cn/papers/126077

---

> Reangle-A-Video：用单一视角重构多彩世界的4D视频生成技术

## 摘要

本文提出了一种名为Reangle-A-Video的框架，用于从单个输入视频生成同步的多视角视频。与主流方法依赖大规模4D数据集训练多视角视频扩散模型不同，该方法将多视角视频生成任务重新定义为视频到视频的翻译问题，并利用公开可用的图像和视频扩散先验。Reangle-A-Video通过两个阶段实现：1) 多视角运动学习：通过对图像到视频扩散Transformer进行自监督微调，提取一组扭曲视频中的视角不变运动；2) 多视角一致的图像到图像翻译：在推理时使用DUSt3R进行跨视角一致性引导，将输入视频的第一帧扭曲并修复成各种相机视角下的起始图像。实验结果表明，Reangle-A-Video在静态视图传输和动态摄像机控制任务中超越了现有方法，为多视角视频生成提供了一种新解决方案。此外，作者计划公开代码和数据。

## Abstract

We introduce Reangle-A-Video, a unified framework for generating synchronized multi-view videos from a single input video. Unlike mainstream approaches that train multi-view video diffusion models on large-scale 4D datasets, our method reframes the multi-view video generation task as video-to-videos translation, leveraging publicly available image and video diffusion priors. In essence, Reangle-A-Video operates in two stages. (1) Multi-View Motion Learning: An image-to-video diffusion transformer is synchronously fine-tuned in a self-supervised manner to distill view-invariant motion from a set of warped videos. (2) Multi-View Consistent Image-to-Images Translation: The first frame of the input video is warped and inpainted into various camera perspectives under an inference-time cross-view consistency guidance using DUSt3R, generating multi-view consistent starting images. Extensive experiments on static view transport and dynamic camera control show that Reangle-A-Video surpasses existing methods, establishing a new solution for multi-view video generation. We will publicly release our code and data. Project page: this https URL
