---
title: PixWorld
canonical_url: "https://www.modelscope.cn/models/SensenGao/PixWorld"
md_url: "https://www.modelscope.cn/models/SensenGao/PixWorld.md"
repository: SensenGao/PixWorld
last_updated: 2026-08-23
base_model_relation: finetune
library_name:
  - safetensors
  - pytorch
frameworks:
  - pytorch
downloads: 253
stars: 0
---

# PixWorld

> PixWorld - SensenGao 在 ModelScope 开源的模型。PixWorld — pixel-space 3D scene generation and reconstruction

- **Repository**: SensenGao/PixWorld
- **Downloads**: 253
- **Stars**: 0
- **Last updated**: 2026-08-23

Source: https://www.modelscope.cn/models/SensenGao/PixWorld

---

# PixWorld — pixel-space 3D scene generation and reconstruction

Two checkpoints of a single model: a diffusion transformer that runs **directly on pixels**
— no VAE, no latent space — and lifts posed multi-view RGB to a **pixel-aligned 3D Gaussian
field** rendered with gsplat.

| | steps | what it is |
|---|---|---|
| `PixWorld-L2P-Wan5B` | 50 | the base model |
| `PixWorld-L2P-Wan5B-4steps` | 4 | distilled from it, for fast sampling |

Both are 5.12B parameters, fine-tuned from
[Wan2.2-TI2V-5B](https://www.modelscope.cn/models/Wan-AI/Wan2.2-TI2V-5B-Diffusers), and both
work at 480×832 over 16 views.

The same weights do three things:

- **text → 3D** — a prompt and a camera path
- **image → 3D** — one photograph, extended to a full multi-view scene
- **reconstruction** — real posed views in, Gaussians out, in a *single* forward pass at
  σ = 0; no sampler and no step count are involved there

## Download

```python
from modelscope import snapshot_download
model_dir = snapshot_download('SensenGao/PixWorld')
```

```
PixWorld-L2P-Wan5B/            PixWorld-L2P-Wan5B.safetensors        + config.json
PixWorld-L2P-Wan5B-4steps/     PixWorld-L2P-Wan5B-4steps.safetensors + config.json
```

## Use

Inference and training code: **https://github.com/SensenGao/PixWorld**

Every run takes an explicit **camera path**, plus a **text prompt**, a **reference image**,
or both. The repo ships real camera paths and reference images under `infer/examples/`.

**text → 3D** — prompt + camera path:

```bash
python infer/infer.py \
    --ckpt <model_dir>/PixWorld-L2P-Wan5B-4steps/PixWorld-L2P-Wan5B-4steps.safetensors \
    --wan_path <Wan2.2-TI2V-5B-Diffusers> \
    --cameras infer/examples/poses/t2mv_living_room.json \
    --prompt "a cozy living room with a stone fireplace and a leather sofa" \
    --out out/living_room
```

**image → 3D** — reference image + camera path (the image is view 0 of that path):

```bash
python infer/infer.py \
    --ckpt <model_dir>/PixWorld-L2P-Wan5B-4steps/PixWorld-L2P-Wan5B-4steps.safetensors \
    --wan_path <Wan2.2-TI2V-5B-Diffusers> \
    --cameras infer/examples/poses/i2mv_bedroom.json \
    --image infer/examples/i2mv_bedroom/reference.jpg \
    --prompt "a bedroom with light walls and a black metal bed" \
    --out out/bedroom
```

`--cameras` takes a JSON of 16 cameras as `[qw qx qy qz | tx ty tz | fx fy cx cy]`, rooted
at view 0 and scaled so the largest translation is 1. Use `--trajectory {dolly,orbit,pan,spiral}`
instead if you want a synthetic path, but a real camera path is what the model was trained on.

The Wan2.2 checkout is needed for its UMT5 text encoder; the transformer weights here are
complete on their own.

Each run writes the generated views, a video rendered from the Gaussian field along the
camera path (`--video_frames 81`, `--video_round_trip` to fly out and back), and `scene.ply`,
which opens in any 3D Gaussian Splatting viewer.

## config.json

These files hold **only** what changes the model's output. There is no training metadata in
them and none in the weights.

| field | |
|---|---|
| `gs_depth_max` | upper bound on the Gaussian head's depth activation, applied as `exp(min(logit, log 50))`. Part of the geometry, not a preference. |
| `n_gen_steps`, `gen_shift`, `gs_step` | *4-step model only* — the fixed sampler ladder it was distilled to, σ = 1.0 / 0.9 / 0.75 / 0.5, with the 3D lift on the last step. Changing them does not give a faster model, it gives a wrong one. |

A checkpoint carrying `n_gen_steps` is a few-step model; one without it is the base model.

The 4-step model samples at **guidance 1.0** — distillation matched it to a guided teacher,
so guidance is already in its weights and applying it again double-counts. `infer.py` reads
this from `config.json` and sets it for you.

Derived from Wan2.2-TI2V-5B; please observe that model's license terms.
