---
title: RynnWorld-Teleop
canonical_url: "https://www.modelscope.cn/models/DAMO_Academy/RynnWorld-Teleop"
md_url: "https://www.modelscope.cn/models/DAMO_Academy/RynnWorld-Teleop.md"
repository: DAMO_Academy/RynnWorld-Teleop
last_updated: 2026-07-08
license: apache-2.0
pipeline_tag: image-to-video
tasks:
  - image-to-video
base_model:
  - Wan-AI/Wan2.2-TI2V-5B-Diffusers
base_model_relation: finetune
library_name:
  - pytorch
frameworks:
  - pytorch
downloads: 5
stars: 1
tags:
  - world-model
  - robotics
  - teleoperation
  - video-generation
  - image-to-video
  - egocentric
  - diffusion
---

# RynnWorld-Teleop

> RynnWorld-Teleop - DAMO_Academy 在 ModelScope 开源的模型。RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation

DAMO_Academy/RynnWorld-Teleop 是 ModelScope 魔搭社区上的image-to-video模型，采用 apache-2.0 许可，基于 Wan-AI/Wan2.2-TI2V-5B-Diffusers 构建。

- **Repository**: DAMO_Academy/RynnWorld-Teleop
- **License**: apache-2.0
- **Tasks**: image-to-video
- **Base model**: Wan-AI/Wan2.2-TI2V-5B-Diffusers
- **Tags**: world-model, robotics, teleoperation, video-generation, image-to-video, egocentric, diffusion
- **Downloads**: 5
- **Stars**: 1
- **Last updated**: 2026-07-08

Source: https://www.modelscope.cn/models/DAMO_Academy/RynnWorld-Teleop

---

<div align="center">

## RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation

</div>


<p align="center">
       💫 <a href="https://alibaba-damo-academy.github.io/RynnWorld-Teleop.github.io/"><b>Project Page</b></a>&nbsp;&nbsp; | &nbsp;&nbsp; 🤗 <a href ="https://huggingface.co/collections/Alibaba-DAMO-Academy/RynnWorld-Teleop"><b> Hugging Face </b></a> &nbsp;&nbsp; | &nbsp;&nbsp; 🤖 <a href = "https://www.modelscope.cn/collections/DAMO_Academy/RynnWorld-Teleop"><b> ModelScope</b></a>  &nbsp;|&nbsp; 🚀 <a href="https://huggingface.co/spaces/Alibaba-DAMO-Academy/RynnWorld-Teleop"><b>Video</b></a> &nbsp;&nbsp; | &nbsp;&nbsp; 📄 <a href="https://arxiv.org/abs/2602.14979v1">arXiv</a>&nbsp;&nbsp;

</p>

---

## 🌟 Abstract

We introduce **RynnWorld-Teleop**, a robot-centric generative world model that instantiates the paradigm of **digital teleoperation**—decoupling robot data collection from physical hardware constraints. By transforming an operator's real-time hand-pose stream into high-fidelity egocentric robotic videos from a single reference image, RynnWorld-Teleop enables the scaling of expert trajectories in a purely virtual environment. Our framework integrates depth-aware skeletal conditioning with a progressive human-to-robot training curriculum, allowing it to inherit rich manipulation priors from large-scale human datasets. To support interactive use, we distill the model into a causal, autoregressive student capable of real-time streaming. Policies trained exclusively on RynnWorld-Teleop synthetic data achieve effective zero-shot Sim2Real transfer, demonstrating its power as a high-fidelity data engine for scaling dexterous robotic learning.

---

## 📰 News
* **[2026.07.07]**  🔥🔥 Release our <a href="https://alibaba-damo-academy.github.io/RynnWorld-Teleop.github.io/assets/RynnWorld-Teleop_Report.pdf">Technical Report</a> !!
* **[2026.07.07]**  🔥🔥 Release our code and model checkpoints!!

---

## 📦 This Repository

This repository hosts the **SFT (full fine-tune) checkpoint** of RynnWorld-Teleop. Given a first-frame image and a hand-pose / skeleton control video, the model generates a high-fidelity egocentric robotic video.

### Model Zoo

| Model            | HuggingFace | ModelScope |
| :--------------- | :---------: | :--------: |
| SFT  | [Link](https://huggingface.co/Alibaba-DAMO-Academy/RynnWorld-Teleop)    | [Link](https://www.modelscope.cn/models/DAMO_Academy/RynnWorld-Teleop)   |
| Causal  | [Link](https://huggingface.co/Alibaba-DAMO-Academy/RynnWorld-Teleop-Causal)    | [Link](https://www.modelscope.cn/models/DAMO_Academy/RynnWorld-Teleop-Causal)   |

---

## 🚀 Quick Start

Please refer to the [code repository](https://github.com/alibaba-damo-academy/RynnWorld-Teleop) for the full setup, training, and inference pipeline.

### 🔧 Environment Setup

```bash
conda create -n "rynnworld-teleop" python=3.10 -y
conda activate rynnworld-teleop
pip3 install torch torchvision --index-url https://download.pytorch.org/whl/cu121
pip install -r requirements.txt
```

### 📖 Download the Checkpoint

Our model is developed on top of [Wan2.2-TI2V-5B](https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B-Diffusers). Download the base model and place it under `pretrained/`:
```
RynnWorld-Teleop/
└── pretrained/
    └── Wan2.2-TI2V-5B-Diffusers/
        ├── model_index.json
        ├── scheduler/
        ├── transformer/
        ├── vae/
        └── ...
```

Then download our fine-tuned weights:
```bash
mkdir -p pretrained/RynnWorld-Teleop
huggingface-cli download Alibaba-DAMO-Academy/RynnWorld-Teleop --local-dir pretrained/RynnWorld-Teleop
```

---

## 🎬 Inference

Given a first-frame image and a control video (hand-pose / OpenPose mp4), the model generates the corresponding egocentric video.

```bash
python inference_user.py \
  --image <first_frame.png> \
  --control_video <control.mp4> \
  --output results/my_demo \
  --prompt "Describe the action in one sentence." \
  --checkpoint <sft_checkpoint_dir> \
  --mode sft \
  --control_type add \
  --seeds "42,123,7"
```

**Required arguments**
- `--image`: first-frame image (jpg/png), automatically resized to 832×480
- `--control_video`: control video mp4 (hand-pose / OpenPose), sampled/interpolated to 81 frames
- `--output`: output directory

**Useful options**
- `--mode sft|lora`: which checkpoint type to load (default: `sft`)
- `--prompt`: optional natural-language description (encoded with the T5 text encoder)
- `--text_embedding`: alternative pre-encoded prompt embedding `.safetensors`
- `--seeds "42,123,7"`: generate multiple samples in one run
- `--no_ema`: use raw weights instead of EMA
- `--guidance_scale`: classifier-free guidance scale (default 1.0)
- `--control_type add|concat|add-plus`: how the control signal is merged

For **real-time streaming inference** with the distilled causal student, please see the [RynnWorld-Teleop-Causal](https://huggingface.co/Alibaba-DAMO-Academy/RynnWorld-Teleop-Causal) checkpoint.

---

## 🏋️ Training

We train the teacher model in **three stages**:

- **Stage 0 — Pretrain (egocentric human videos):** Full-parameter SFT on large-scale egocentric data, no control-video conditioning. Absorbs general manipulation priors.
- **Stage 1 — Control-conditioned fine-tuning:** Adds a zero-initialized `control_patch_embedding` (Conv3d) and a learnable `control_scale` to inject hand-pose control video into the diffusion process. LoRA (lightweight) and Full-SFT (best quality) variants are provided.
- **Distillation → Causal student:** Distilled into a causal autoregressive model for real-time streaming.

Full training scripts, configs, and data-preparation instructions are available in the [code repository](https://github.com/alibaba-damo-academy/RynnWorld-Teleop).

---

## 📑 Citation

If you find this project useful, please cite:

```bibtex
@article{rynnworld_teleop,
  title  = {RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation},
  author = {DAMO Academy, Alibaba Group},
  year   = {2026},
}
```

## License

Apache License 2.0 — see [LICENSE](https://github.com/alibaba-damo-academy/RynnWorld-Teleop/blob/main/LICENSE) for details.
