---
title: MoGe4D
canonical_url: "https://www.modelscope.cn/models/YanranZhang/MoGe4D"
md_url: "https://www.modelscope.cn/models/YanranZhang/MoGe4D.md"
repository: YanranZhang/MoGe4D
last_updated: 2026-07-06
license: apache-2.0
pipeline_tag: image-to-video
tasks:
  - image-to-video
parameters: 17.9B
tensor_type:
  - BF16
library_name:
  - safetensors
  - diffusers
  - pytorch
frameworks:
  - Pytorch
language:
  - en
downloads: 33
stars: 1
tags:
  - 4d-generation
  - image-to-4d
  - diffusion
  - novel-view-synthesis
  - point-trajectory
---

# MoGe4D

> MoGe4D - YanranZhang 在 ModelScope 开源的模型。[ECCV 2026] MoGe4D: Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation

YanranZhang/MoGe4D 是 ModelScope 魔搭社区上的 17.9B 参数image-to-video模型，采用 apache-2.0 许可。

- **Repository**: YanranZhang/MoGe4D
- **License**: apache-2.0
- **Tasks**: image-to-video
- **Parameters**: 17.9B
- **Tags**: 4d-generation, image-to-4d, diffusion, novel-view-synthesis, point-trajectory
- **Downloads**: 33
- **Stars**: 1
- **Last updated**: 2026-07-06

Source: https://www.modelscope.cn/models/YanranZhang/MoGe4D

---

# [ECCV 2026] MoGe4D: Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation

<b>[Yanran Zhang](https://github.com/Zhangyr2022/)<sup>\*,1</sup>, [Ziyi Wang](https://wangzy22.github.io/)<sup>\*,1</sup>, [Wenzhao Zheng](https://wzzheng.net/#)<sup>†,1</sup>, [Zheng Zhu](http://www.zhengzhu.net/)<sup>2</sup>, [Jie Zhou](https://scholar.google.com/citations?user=6a79aPwAAAAJ&hl=en)<sup>1</sup>, [Jiwen Lu](https://ivg.au.tsinghua.edu.cn/Jiwen_Lu/)<sup>1</sup></b>

<sup>1</sup>Department of Automation, Tsinghua University &nbsp;&nbsp;&nbsp; <sup>2</sup>GigaAI

<i><sup>*</sup>Equal Contribution &nbsp;&nbsp; <sup>†</sup>Corresponding Author</i>

<p align="center">
  <a href="https://github.com/Zhangyr2022/MoGe4D"><img src="https://img.shields.io/badge/GitHub-Code-black?logo=github" alt="GitHub"></a>
  <a href="https://arxiv.org/abs/2512.05044"><img src="https://img.shields.io/badge/arXiv-Paper-b31b1b?logo=arxiv&logoColor=white" alt="arXiv"></a>
  <a href="https://ivg-yanranzhang.github.io/MoGe4D/"><img src="https://img.shields.io/badge/Project-Website-blue?logo=googlechrome&logoColor=white" alt="Project"></a>
  <br>
  <a href="https://www.modelscope.cn/models/YanranZhang/MoGe4D"><img src="https://img.shields.io/badge/🤖%20ModelScope-Model-4e29ff" alt="ModelScope Model"></a>
  <a href="https://huggingface.co/Yanran21/MoGe4D"><img src="https://img.shields.io/badge/🤗%20HuggingFace-Model-ffd21e" alt="HuggingFace Model"></a>
  <a href="https://www.modelscope.cn/datasets/YanranZhang/TrajScene-60K"><img src="https://img.shields.io/badge/🤖%20ModelScope-Dataset-4e29ff" alt="ModelScope Dataset"></a>
</p>

## 📄 Paper Summary

Generating interactive and dynamic 4D scenes from a single static image is a core challenge. Existing methods decouple geometry from motion — either *generate-then-reconstruct* (geometric inconsistency) or *reconstruct-then-generate* (limited, externally-constrained motion) — causing spatiotemporal inconsistency and poor generalization.

**MoGe4D** (Motion and Geometry-aware image-to-4D synthesis) is a geometry-conditioned framework that models a scene as **dense 4D point trajectories**. Starting from an initial geometric prior of the input image, it predicts future time-varying trajectories through a diffusion process, tightly coupling geometric modeling with motion generation. This yields 4D scenes with strong temporal coherence, geometry-aware consistency, and compelling novel-view synthesis.

**Contributions:**
- **TrajScene-60K** — 60K videos with dense 4D point trajectories (3M+ frames, ~12B 3D points).
- **4D-STraG** — a diffusion trajectory generator with *depth-guided motion normalization* and a *Motion Perception Module (MPM)*.
- **4D-ViSM** — a view-synthesis module rendering the 4D representation under arbitrary camera trajectories.

## 🧱 Model Structure

This repository releases the three trained components of MoGe4D:

| Path | Size | Description |
|---|---|---|
| `4D-STraG/diffusion_pytorch_model.safetensors` | ~31.9 GiB | 4D Scene Trajectory Generator (diffusion model, built on Wan2.1-14B) |
| `4D-ViSM/lora_diffusion_pytorch_model.safetensors` | ~1.36 GiB | 4D View Synthesis Module (LoRA adapter) |
| `VAE/vae/pytorch_model.bin` | ~484 MiB | Motion-sensitive VAE for trajectory signals |
| `VAE/{encoder,decoder}_prompt/pytorch_model.bin` | ~1–2 MiB | VAE prompt encoder/decoder |
| `VAE/{optimizer.bin, scheduler.bin, random_states_0.pkl}` | — | Training states for the VAE (optional, for resuming training) |

## 🛠️ Usage

### 1. Set up the environment

```bash
git clone https://github.com/Zhangyr2022/MoGe4D.git
cd MoGe4D
conda create -n MoGe4D python=3.10 && conda activate MoGe4D
conda install pytorch torchvision torchaudio pytorch-cuda=12.4 -c pytorch -c nvidia
pip install -r requirements.txt
```

Install third-party deps: [UniDepth](https://github.com/lpiccinelli-eth/UniDepth) and [diff-gaussian-rasterization](https://github.com/graphdeco-inria/diff-gaussian-rasterization).

### 2. Download the checkpoints

```bash
huggingface-cli download Yanran21/MoGe4D --local-dir ./models --resume-download
```

(Also place the base backbones [Wan2.1-Fun-V1.1-14B-Control/InP](https://huggingface.co/alibaba-pai), [OmniMAE](https://dl.fbaipublicfiles.com/omnivore/omnimae_ckpts/vitb_pretrain.torch), and [UniDepth](https://huggingface.co/lpiccinelli/unidepth-v2-vitl14) under `./models`.)

### 3. Inference

```bash
bash scripts/inference/infer.sh        # whole pipeline: image → 4D scene → multi-view videos
```

See the [GitHub README](https://github.com/Zhangyr2022/MoGe4D) for training scripts and details.

## 📊 Results

MoGe4D delivers superior geometric consistency, dynamic realism, and visual fidelity over decoupled approaches (e.g., generate-then-reconstruct with VGGT). Please refer to the paper for quantitative metrics and qualitative comparisons.

## 📖 Citation

```bibtex
@inproceedings{zhang2026moge4d,
  title={Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation},
  author={Zhang, Yanran and Wang, Ziyi and Zheng, Wenzhao and Zhu, Zheng and Zhou, Jie and Lu, Jiwen},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2026}
}
```

## 📧 Contact

- Yanran Zhang — zhangyr21@mails.tsinghua.edu.cn


```

## 📧 Contact

- Yanran Zhang — zhangyr21@mails.tsinghua.edu.cn
