---
title: VPP2
canonical_url: "https://www.modelscope.cn/models/haodong123/VPP2"
md_url: "https://www.modelscope.cn/models/haodong123/VPP2.md"
repository: haodong123/VPP2
chinese_name: "Video Prediction Policy 2"
last_updated: 2026-10-10
model_type:
  - vpp2
base_model:
  - Wan-AI/Wan2.1-I2V-14B-480P
base_model_relation: finetune
library_name:
  - pytorch
frameworks:
  - Pytorch
language:
  - en
downloads: 2
stars: 0
tags:
  - robotics
  - video-prediction
  - robodojo
  - libero
  - "arxiv:2610.10270"
---

# VPP2

> VPP2 - haodong123 在 ModelScope 开源的模型。Video Prediction Policy 2: Predict Better, Act Better

haodong123/VPP2 是 ModelScope 魔搭社区上的机器学习模型，基于 Wan-AI/Wan2.1-I2V-14B-480P 构建。

- **Repository**: haodong123/VPP2
- **Base model**: Wan-AI/Wan2.1-I2V-14B-480P
- **Tags**: robotics, video-prediction, robodojo, libero, arxiv:2610.10270
- **Downloads**: 2
- **Stars**: 0
- **Last updated**: 2026-10-10

Source: https://www.modelscope.cn/models/haodong123/VPP2

---

# Video Prediction Policy 2: Predict Better, Act Better

Official VPP2 weights for zero-shot video prediction, RoboDojo and LIBERO.

[Paper](https://arxiv.org/abs/2610.10270) · [Code](https://github.com/roboterax/video-prediction-policy-2) · [Project page](https://robert-gyj.github.io/video-prediction-policy-2/) · [Hugging Face](https://huggingface.co/Haodong082399/VPP2)

## Model weights

| Checkpoint | Path | Use |
|---|---|---|
| VPP2 Stage-1 Video (49 frames) | `checkpoints_video/vpp2-video-stage1-49f.pth` | Zero-shot image-to-video prediction |
| VPP2 Stage-2 Video (17 frames) | `checkpoints_video/vpp2-video-stage2-17f.pth` | Zero-shot image-to-video prediction |
| RoboDojo history-conditioned Video-10k | `checkpoints/initialization/robodojo_his10k.pt` | Joint + Action2B training initializer |
| RoboDojo joint Video-100k | `checkpoints/joint2b_s100000/video.pt` | Paired RoboDojo evaluation |
| RoboDojo Action2B-100k | `checkpoints/joint2b_s100000/action.pt` | Paired RoboDojo evaluation |
| LIBERO Video-10k | `checkpoints/libero/video_step010000.pt` | Action training and evaluation |
| LIBERO Action-30k | `checkpoints/libero/action_step030000.pt` | LIBERO, LIBERO-OOD and LIBERO-PRO evaluation |

Shared VAE, CLIP, UMT5 and tokenizer assets are under
`checkpoints/Wan2.1-I2V-14B-480P/`. Each policy pair includes its manifest and
normalization statistics. Keep each pair together.

## Download

Install the ModelScope CLI:

```bash
python -m pip install -U modelscope
```

Each command downloads the root [`config.json`](config.json), which lists the
released checkpoints, video frame counts and inference defaults, together with
the selected weights and shared Wan encoders.

### Stage-1 video

Use `--num-frames 49`.

```bash
modelscope download --model haodong123/VPP2 --local_dir weights \
  --include 'config.json' \
            'checkpoints_video/vpp2-video-stage1-49f.pth' \
            'checkpoints/Wan2.1-I2V-14B-480P/**'
```

### Stage-2 video

Use `--num-frames 17`.

```bash
modelscope download --model haodong123/VPP2 --local_dir weights \
  --include 'config.json' \
            'checkpoints_video/vpp2-video-stage2-17f.pth' \
            'checkpoints/Wan2.1-I2V-14B-480P/**'
```

### RoboDojo training

Use Video-10k to initialize joint Video + Action2B training.

```bash
modelscope download --model haodong123/VPP2 --local_dir weights \
  --include 'config.json' \
            'checkpoints/initialization/robodojo_his10k.pt' \
            'checkpoints/Wan2.1-I2V-14B-480P/**'
```

### RoboDojo evaluation

Download the paired Video-100k and Action2B-100k weights with their normalization statistics.

```bash
modelscope download --model haodong123/VPP2 --local_dir weights \
  --include 'config.json' \
            'checkpoints/joint2b_s100000/**' \
            'checkpoints/Wan2.1-I2V-14B-480P/**'
```

### LIBERO

Use the same Video-10k and Action-30k pair for LIBERO, LIBERO-OOD and LIBERO-PRO.

```bash
modelscope download --model haodong123/VPP2 --local_dir weights \
  --include 'config.json' \
            'checkpoints/libero/**' \
            'checkpoints/Wan2.1-I2V-14B-480P/**'
```

Both standalone video checkpoints use 30 denoising steps, CFG=4 and sigma shift 3.
Follow the [video prediction guide](https://github.com/roboterax/video-prediction-policy-2/blob/main/docs/video_prediction.md)
for image preparation and inference, or the
[RoboDojo](https://github.com/roboterax/video-prediction-policy-2/blob/main/docs/robodojo.md)
and [LIBERO](https://github.com/roboterax/video-prediction-policy-2/blob/main/docs/libero.md)
guides for policy training and evaluation.

### All weights

To download the complete repository, including `config.json`:

```bash
modelscope download --model haodong123/VPP2 --local_dir weights
```

## Evaluation settings

| Benchmark | Checkpoint pair | Denoising | Action horizon / execution |
|---|---|---|---|
| RoboDojo | Video-100k + Action2B-100k | 10 Euler steps, shift 1 | 32 / 24 |
| LIBERO / OOD / PRO | Video-10k + Action-30k | 10 Euler steps, shift 5 | 32 / 10 |

RoboDojo uses its EE16 z-score statistics; LIBERO uses its own 7D action min/max
statistics. Use each benchmark's configuration and simulator setup.

## RoboDojo adapter

The standalone adapter imports `vpp2` from the bundled `VPP2/` directory.
Download it into an existing RoboDojo installation:

```bash
cd /path/to/RoboDojo
modelscope download --model haodong123/VPP2 \
  --local_dir XPolicyLab/policy/VPP2 \
  --include 'config.json' '__init__.py' 'model.py' 'deploy.py' 'launch_policy.py' \
            'eval.sh' 'install.sh' 'setup_eval_policy_server.sh' 'setup_eval_env_client.sh' \
            'deploy.yml' 'README.md' 'VPP2/**' \
            'checkpoints/joint2b_s100000/**' 'checkpoints/Wan2.1-I2V-14B-480P/**'
cd XPolicyLab/policy/VPP2
bash install.sh vpp2
bash eval.sh RoboDojo cover_blocks joint2b_s100000 arx_x5 ee 1 0 1 vpp2 robodojo_env
```

Replace `robodojo_env` with the simulator environment. The policy and simulator
run in separate environments. Full training, data preparation and evaluation
instructions are in the [official code repository](https://github.com/roboterax/video-prediction-policy-2).

## Citation

See [CITATION.bib](CITATION.bib) for the paper citation.

## Licenses

Shared Wan assets retain their upstream license in
[`checkpoints/Wan2.1-I2V-14B-480P/LICENSE.txt`](checkpoints/Wan2.1-I2V-14B-480P/LICENSE.txt).
The official implementation has its own [code license](https://github.com/roboterax/video-prediction-policy-2/blob/main/LICENSE).
