---
title: RynnWorld-4D
canonical_url: "https://www.modelscope.cn/models/DAMO_Academy/RynnWorld-4D"
md_url: "https://www.modelscope.cn/models/DAMO_Academy/RynnWorld-4D.md"
repository: DAMO_Academy/RynnWorld-4D
last_updated: 2026-07-06
license: apache-2.0
pipeline_tag: image-to-video
tasks:
  - image-to-video
base_model:
  - Wan-AI/Wan2.2-TI2V-5B-Diffusers
base_model_relation: finetune
library_name:
  - pytorch
frameworks:
  - Pytorch
downloads: 16
stars: 2
tags:
  - world-model
  - robotics
  - video-generation
  - diffusion
  - 4d
  - rgb-depth-flow
  - embodied-ai
  - manipulation
---

# RynnWorld-4D

> RynnWorld-4D - DAMO_Academy 在 ModelScope 开源的模型。4D Embodied World Models for Robotic Manipulation

DAMO_Academy/RynnWorld-4D 是 ModelScope 魔搭社区上的image-to-video模型，采用 apache-2.0 许可，基于 Wan-AI/Wan2.2-TI2V-5B-Diffusers 构建。

- **Repository**: DAMO_Academy/RynnWorld-4D
- **License**: apache-2.0
- **Tasks**: image-to-video
- **Base model**: Wan-AI/Wan2.2-TI2V-5B-Diffusers
- **Tags**: world-model, robotics, video-generation, diffusion, 4d, rgb-depth-flow, embodied-ai, manipulation
- **Downloads**: 16
- **Stars**: 2
- **Last updated**: 2026-07-06

Source: https://www.modelscope.cn/models/DAMO_Academy/RynnWorld-4D

---

# RynnWorld-4D

**4D Embodied World Models for Robotic Manipulation**

<p align="center">
  💫 <a href="https://alibaba-damo-academy.github.io/RynnWorld-4D.github.io/"><b>Project Page</b></a>&nbsp;&nbsp;|&nbsp;&nbsp;
  🤗 <a href="https://huggingface.co/collections/Alibaba-DAMO-Academy/RynnWorld-4D"><b>Hugging Face Collection</b></a>&nbsp;&nbsp;|&nbsp;&nbsp;
  🤖 <a href="https://www.modelscope.cn/collections/DAMO_Academy/RynnWorld-4D"><b>ModelScope</b></a>&nbsp;&nbsp;|&nbsp;&nbsp;
  📄 <a href="https://alibaba-damo-academy.github.io/RynnWorld-4D.github.io/assets/RynnWorld-4D_Report.pdf"><b>Technical Report</b></a>
</p>

---

## Model Summary

**RynnWorld-4D** is a 4D embodied world model that generates synchronized **RGB, depth, and optical flow** (RGB-DF) videos from a single reference image and text prompt. Unlike conventional 2D pixel-prediction world models, RynnWorld-4D captures the underlying 3D geometry and temporal motion trajectories of a scene, producing a physically grounded representation that bridges generative world modeling and low-level robotic control.

This repository hosts the **Stage-3 checkpoint** (tri-branch full-parameter SFT with joint cross-modal attention) built on top of [Wan2.2-TI2V-5B-Diffusers](https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B-Diffusers).

### Key Highlights

- **Projective 4D Representation** — Unified RGB-DF format lets pixels be unprojected into metric 3D scene flow for precise geometric and kinetic grounding.
- **Tri-Branch Diffusion Architecture** — Three dedicated transformer branches (RGB / depth / flow) with mutual cross-modal joint attention ensure appearance, geometry, and motion evolve with high spatio-temporal consistency.
- **Cosine-Decay Joint Injection** — During Stage-3 training, the depth/flow → RGB injection is gradually annealed from 1.0 to 0.0, preserving cross-modal consistency while preventing RGB quality degradation.
- **Action-from-Latent Policy (companion model)** — The paired [RynnWorld-4D-Policy](https://alibaba-damo-academy.github.io/RynnWorld-4D.github.io/) consumes internal 4D latents directly for high-frequency (9 Hz+), closed-loop bimanual manipulation.

---

## Model Architecture

- **Backbone:** Wan2.2-TI2V-5B (Diffusion Transformer, ~5B params)
- **Branches:** 3 parallel transformer streams — RGB, Depth, Optical Flow
- **Cross-modal fusion:** Joint attention with 3D RoPE, applied every 3 layers across all 30 blocks
  - `--joint_start_layer 0 --joint_end_layer 30 --joint_every_n_layers 3`
  - `--joint_frame_wise True` (attention restricted to same-frame tokens)
  - `--joint_use_rope True`
- **Fusion mode:** bidirectional joint attention with cosine-decayed video injection

---

## Files

```
RynnWorld-4D/
├── pytorch_model/
│   └── mp_rank_00_model_states.pt   # 28 GB — full trained weights
└── ema_weights.pt                   # 28 GB — EMA shadow weights (recommended for inference)
```

Both files together (~56 GB) are required for the standard inference path in `inference-sft.py`. The EMA weights are loaded on top of the base weights and typically yield the best sample quality.

---

## Quick Start

### 1. Environment

```bash
git clone https://github.com/Alibaba-DAMO-Academy/RynnWorld-4D
cd RynnWorld-4D

conda create -n rynnworld4d python=3.10 -y
conda activate rynnworld4d
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu121
pip install -r requirements.txt
pip install -e . --no-build-isolation
```

### 2. Download the base backbone and this checkpoint

```bash
# base backbone (required)
huggingface-cli download Wan-AI/Wan2.2-TI2V-5B-Diffusers \
    --local-dir ./pretrained/Wan2.2-TI2V-5B-Diffusers

# RynnWorld-4D Stage-3 weights (this repo)
huggingface-cli download Alibaba-DAMO-Academy/RynnWorld-4D \
    --local-dir ./pretrained/RynnWorld-4D
```

### 3. Inference

```bash
python inference-sft.py \
    --model_path       ./pretrained/Wan2.2-TI2V-5B-Diffusers \
    --checkpoint_path  ./pretrained/RynnWorld-4D \
    --json_path        ./data/sample.json \
    --output_dir       ./results/rynnworld4d \
    --fusion_mode      joint \
    --share_ffn        False \
    --joint_start_layer 0 \
    --joint_end_layer   30 \
    --joint_every_n_layers 3 \
    --joint_frame_wise  True \
    --joint_use_rope    True \
    --joint_unidirectional False \
    --zero_fusion       False \
    --use_ema           True \
    --num_inference_steps 50 \
    --guidance_scale    1.0
```

**Config flags must match training.** The important ones for this checkpoint:

| Flag | Value | Reason |
|---|---|---|
| `--fusion_mode` | `joint` | Stage-3 uses joint cross-modal attention. |
| `--joint_use_rope` | `True` | 3D RoPE is enabled in joint attention. |
| `--joint_unidirectional` | `False` | Trained as bidirectional (with video-side cosine decay). |
| `--use_ema` | `True` | Loads EMA weights on top of base for best quality. |
| `--zero_fusion` | `False` | Fusion weights are trained; do not zero them. |

Each sample produces three synchronized streams:

```
<output_dir>/<sample_id>/
├── rgb.mp4       # generated RGB
├── depth.mp4     # generated depth
└── flow.mp4      # generated optical flow
```

---

## Training Recipe (Summary)

RynnWorld-4D is trained in **three stages** on top of the Wan2.2-TI2V-5B backbone. This checkpoint is the output of Stage 3.

| Stage | Objective | Trainable | Fusion |
|---|---|---|---|
| **1** | Full-parameter SFT — warm up all three branches (RGB / depth / flow) independently. | All branches | None |
| **2** | Enable 3D-RoPE joint cross-modal attention; freeze non-joint params. Train bidirectional joint layers. | Joint attention only | Bidirectional, RoPE |
| **3** | Full-parameter fine-tuning with **cosine-decay video injection**: the depth/flow → RGB gate multiplier anneals from `1.0 → 0.0` over `--joint_video_decay_steps`, letting the model absorb early-stage cross-modal consistency while gradually restoring independent RGB generation. | Everything | Bidirectional → unidirectional (decayed) |

See the training scripts (`scripts/rynnworld4d-stage{1,2,3}.sh`) in the code repository for full hyperparameters.

---

## Intended Uses

- **Research** on 4D world models, cross-modal video generation, and geometry-aware embodied AI.
- **Feature extraction** for downstream robotic policy learning (see `RynnWorld-4D-Policy`).
- **Depth / flow-consistent video synthesis** conditioned on a first-frame image + text prompt.

### Out-of-Scope

- Non-embodied / non-robotic content generation is possible but not the training focus.
- The model is not suitable for real-time / on-device inference without further distillation.
- Not designed for photorealistic humans, faces, or open-domain creative video generation.

---

## Limitations

- Trained primarily on robotic manipulation and egocentric datasets; performance on unrelated domains (natural scenes, cinematic footage) may be limited.
- 25-frame, 480×832 latent resolution — extended sequences may drift.
- Bidirectional joint attention can occasionally propagate depth/flow artifacts to RGB, mitigated but not eliminated by the cosine-decay schedule.
- Guidance-scale > 1.0 requires a null-prompt embedding file (see `inference-sft.py`).

---

## Companion Model

- **[RynnWorld-4D-Policy](https://huggingface.co/collections/Alibaba-DAMO-Academy/RynnWorld-4D)** — a lightweight flow-matching action head trained on top of the frozen RynnWorld-4D backbone for high-frequency bimanual manipulation.

---

## Acknowledgements

Built on top of and inspired by [Wan2.2-TI2V-5B](https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B-Diffusers), [Depth-Anything-3](https://github.com/ByteDance-Seed/Depth-Anything-3), [Video Prediction Policy (VPP)](https://github.com/roboterax/video-prediction-policy), and [ptlflow](https://github.com/hmorimitsu/ptlflow).

---

## Citation

```bibtex
@article{rynnworld4d,
  title  = {RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation},
  author = {DAMO Academy, Alibaba Group},
  year   = {2026},
}
```

---

## License

Apache License 2.0. Vendored third-party code follows its original upstream licenses (preserved in `third_party/*/LICENSE`).
