---
title: Kai0
canonical_url: "https://www.modelscope.cn/datasets/OpenDriveLab/Kai0"
md_url: "https://www.modelscope.cn/datasets/OpenDriveLab/Kai0.md"
repository: OpenDriveLab/Kai0
last_updated: 2026-06-27
license: cc-by-nc-sa-4.0
storage_size: "360 GB"
downloads: 1586232
stars: 1
---

# Kai0

> Kai0 - OpenDriveLab 在 ModelScope 开源的数据集。TODO [ ] The advantage label will be coming soon.

OpenDriveLab/Kai0 是 ModelScope 魔搭社区上的数据集，存储大小 360 GB，采用 cc-by-nc-sa-4.0 许可。

- **Repository**: OpenDriveLab/Kai0
- **License**: cc-by-nc-sa-4.0
- **Storage size**: 360 GB
- **Downloads**: 1586232
- **Stars**: 1
- **Last updated**: 2026-06-27

Source: https://www.modelscope.cn/datasets/OpenDriveLab/Kai0

---

# KAI0
<div align="center">
  <a href="">
    <img src="https://img.shields.io/badge/GitHub-grey?logo=GitHub" alt="GitHub Badge">
  </a>
  <a href="https://www.modelscope.cn/models/OpenDriveLab/Kai0">
    <img src="https://img.shields.io/badge/Model-grey?logo=Model" alt="Model Badge">
  </a>
  <a href="https://mmlab.hk/research/kai0">
    <img src="https://img.shields.io/badge/Research_Blog-grey?style=flat" alt="Research Blog Badge">
  </a>
</div>

# TODO
- [ ] The advantage label will be coming soon.

## Contents
- [About the Dataset](#about-the-dataset)
- [Load the Dataset](#get-started)
- [Download the Dataset](#download-the-dataset)
- [Dataset Structure](#dataset-structure)
    - [Folder hierarchy](#folder-hierarchy)
    - [Details](#details)
- [License and Citation](#license-and-citation)

## [About the Dataset](#contents)
- **~134 hours** real world scenarios 
- **Main Tasks**
    - ***Task_A***
      - Single task
      - Initial state: T-shirts are randomly tossed onto the table, presenting random crumpled configurations
      - Manipulation task: Operate the robotic arm to unfold the garment, then fold it
    - ***Task_B***
      - Garment classification and arrangement task
      - Initial state: Randomly pick a garment from the laundry basket
      - Classification: Determine whether the garment is a T-shirt or a dress shirt
      - Manipulation task:
        - If it is a T-shirt, fold the garment
        - If it is a dress shirt, expose the collar, then push it to one side of the table
    - ***Task_C***
      - Single task
      - Initial state: Hanger is randomly placed, garment is randomly positioned on the table
      - Manipulation task: Operate the robotic arm to thread the hanger through the garment, then hang it on the rod
- **Count of the dataset** 

    | Task | Base (episodes count/hours) | DAgger (episodes count/hours) | Total(episodes count/hours) |
    |------|-----------------------------|-------------------------------|-----------------------------|
    | Task_A | 3,055/~42 hours             | 3,457/ ~13 Hours              | 6,512 /~55 hours            |
    | Task_B | 5988/~31 hours              | 769/~22 hours                 | 6757/~53 hours              |
    | Task_C | 6954/~61 hours              | 686/~12 hours                 | 7640/~73 hours              |
    | **Total** | **15,997/~134 hours**       | **4,912/~47 hours**           | **20,909/~181 hours**       |

## [Load the dataset](#contents)
- This dataset was created using [LeRobot](https://github.com/huggingface/lerobot)
- The dataset's version is LeRobotDataset v2.1
### For LeRobot version < 0.4.0
Choose the appropriate import based on your version:

| Version                | Import Path |
|------------------------|-------------|
| `<= 0.1.0`             | `from lerobot.common.datasets.lerobot_dataset import LeRobotDataset` |
| `> 0.1.0` and `< 0.4.0` | `from lerobot.datasets.lerobot_dataset import LeRobotDataset` |

```python
# For version <= 0.1.0
from lerobot.common.datasets.lerobot_dataset import LeRobotDataset

# For version > 0.1.0 and < 0.4.0
from lerobot.datasets.lerobot_dataset import LeRobotDataset

# Load the dataset
dataset = LeRobotDataset(repo_id='where/the/dataset/you/stored')
```

### For LeRobot version >= 0.4.0

You need to migrate the dataset from v2.1 to v3.0 first. See the official documentation: [Migrate the dataset from v2.1 to v3.0](https://huggingface.co/docs/lerobot/lerobot-dataset-v3)

```bash
python -m lerobot.datasets.v30.convert_dataset_v21_to_v30 --repo-id=<HF_USER/DATASET_ID>
```

## [Download the Dataset](#contents)
### Git 下载
```bash
#Git模型下载
git clone https://www.modelscope.cn/datasets/OpenDriveLab/Kai0.git
```
### Terminal (CLI)
```bash
# Download a single file
modelscope download --dataset OpenDriveLab-org/kai0 README.md \
                    --local_dir "/where/you/want/to/save"
# Download the entire dataset
modelscope download --dataset OpenDriveLab-org/kai0 \
                    --local-dir "/where/you/want/to/save"
```

## [Dataset Structure](#contents)

### [Folder hierarchy](#contents)
Under each task directory, data is partitioned into two subsets: base and dagger.
- base 
  contains 
  original demonstration trajectories of robotic arm manipulation for garment arrangement tasks.
- dagger  
  contains on-policy recovery trajectories collected via iterative DAgger, designed to populate failure recovery modes absent in static demonstrations.
```text
Kai0-data/
├── Task_A/
│   ├── base/
│   │   ├── data/
│   │   │   ├── chunk-000/
│   │   │   │   ├── episode_000000.parquet
│   │   │   │   ├── episode_000001.parquet
│   │   │   │   └── ...
│   │   │   └── ...
│   │   ├── videos/
│   │   │   ├── chunk-000/
│   │   │   │   ├── observation.images.hand_left/
│   │   │   │   │   ├── episode_000000.mp4
│   │   │   │   │   ├── episode_000001.mp4
│   │   │   │   │   └── ...
│   │   │   │   ├── observation.images.hand_right/
│   │   │   │   │   ├── episode_000000.mp4
│   │   │   │   │   ├── episode_000001.mp4
│   │   │   │   │   └── ...
│   │   │   │   ├── observation.images.top_head/
│   │   │   │   │   ├── episode_000000.mp4
│   │   │   │   │   ├── episode_000001.mp4
│   │   │   │   │   └── ...
│   │   │   │   └── ...
│   │   │   └── ...
│   │   └── meta/
│   │       ├── info.json
│   │       ├── episodes.jsonl
│   │       ├── tasks.jsonl
│   │       └── episodes_stats.jsonl
│   └── dagger/
├── Task_B/
│   ├── base/
│   └── dagger/
├── Task_C/
│   ├── base/
│   └── dagger/
└── README.md
```

<a id='Details'></a>
### [Details](#contents)
#### info.json
the basic struct of the [info.json](#meta/info.json)
```json
{
    "codebase_version": "v2.1",
    "robot_type": "agilex",
    "total_episodes": ...,  # the total episodes in the dataset
    "total_frames": ...,    # The total number of video frames in any single camera perspective
    "total_tasks": ...,     # Total number of tasks
    "total_videos": ...,    # The total number of videos from all camera perspectives in the dataset
    "total_chunks": ...,    # The number of chunks in the dataset
    "chunks_size": ...,     # The max number of episodes in a chunk
    "fps": ...,             # Video frame rate per second
    "splits": {             # how to split the dataset
        "train": ...       
    },
    "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
    "video_path": "videos/chunk-{episode_chunk:03d}/{video_key}/episode_{episode_index:06d}.mp4",
    "features": {
        "observation.images.top_head": {   # the camera perspective
            "dtype": "video",
            "shape": [
                480,
                640,
                3
            ],
            "names": [
                "height",
                "width",
                "channel"
            ],
            "info": {
                "video.height": 480,
                "video.width": 640,
                "video.codec": "av1",
                "video.pix_fmt": "yuv420p",
                "video.is_depth_map": false,
                "video.fps": 30,
                "video.channels": 3,
                "has_audio": false
            }
        },
        "observation.images.hand_left": {   # the camera perspective
            ...
        },
        "observation.images.hand_right": {   # the camera perspective
            ...
        },
        "observation.state": {
            "dtype": "float32",
            "shape": [
                14
            ],
            "names": null
        },
        "action": {
            "dtype": "float32",
            "shape": [
                14
            ],
            "names": null
        },
        "timestamp": {
            "dtype": "float32",
            "shape": [
                1
            ],
            "names": null
        },
        "frame_index": {
            "dtype": "int64",
            "shape": [
                1
            ],
            "names": null
        },
        "episode_index": {
            "dtype": "int64",
            "shape": [
                1
            ],
            "names": null
        },
        "index": {
            "dtype": "int64",
            "shape": [
                1
            ],
            "names": null
        },
        "task_index": {
            "dtype": "int64",
            "shape": [
                1
            ],
            "names": null
        }
    }
}
```

#### [Parquet file format](#contents)
| Field Name | shape | Meaning |
|------------|-------------|-------------|
| observation.state | [N, 14] |left `[:, :6]`, right `[:, 7:13]`, joint angle<br> left`[:, 6]`, right `[:, 13]` , gripper open range|
| action | [N, 14]  |left `[:, :6]`, right `[:, 7:13]`, joint angle<br>left`[:, 6]`, right `[:, 13]` , gripper open range |
| timestamp | [N, 1] | Time elapsed since the start of the episode (in seconds) |
| frame_index | [N, 1] | Index of this frame within the current episode (0-indexed) |
| episode_index | [N, 1] | Index of the episode this frame belongs to |
| index | [N, 1] | Global unique index across all frames in the dataset |
| task_index | [N, 1] | Index identifying the task type being performed |

### [tasks.jsonl](#Task_A/meta/tasks.jsonl)
Contains task language prompts (natural language instructions) that specify the manipulation task to be performed. Each entry maps a task_index to its corresponding task description, which can be used for language-conditioned policy training.

# License and Citation
All the data and code within this repo are under [](). Please consider citing our project if it helps your research.

```BibTeX
@misc{,
  title={},
  author={},
  howpublished={\url{}},
  year={}
}
