---
title: RealVLG-11B
canonical_url: "https://www.modelscope.cn/datasets/cslinfeili/RealVLG-11B"
md_url: "https://www.modelscope.cn/datasets/cslinfeili/RealVLG-11B.md"
repository: cslinfeili/RealVLG-11B
last_updated: 2026-05-13
license: "Apache License 2.0"
storage_size: "139 GB"
downloads: 15128
stars: 0
---

# RealVLG-11B

> RealVLG-11B - cslinfeili 在 ModelScope 开源的数据集。RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation [CVPR 2026] Linfei Li · Lin Zhang · Ying Shen

cslinfeili/RealVLG-11B 是 ModelScope 魔搭社区上的数据集，存储大小 139 GB，采用 Apache License 2.0 许可。

- **Repository**: cslinfeili/RealVLG-11B
- **License**: Apache License 2.0
- **Storage size**: 139 GB
- **Downloads**: 15128
- **Stars**: 0
- **Last updated**: 2026-05-13

Source: https://www.modelscope.cn/datasets/cslinfeili/RealVLG-11B

---

<p align="center">
  <h1 align="center">
    RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation
    <br>
    [CVPR 2026]
  </h1>
  <p align="center">
  <a href="https://lif314.github.io/"><strong>Linfei Li</strong></a>
  ·
  <a href="https://scholar.google.com/citations?user=8VOk_S4AAAAJ&hl=en"><strong>Lin Zhang*</strong></a>
  ·
  <a href="https://scholar.google.com/citations?user=A0N_mS0AAAAJ&hl=en"><strong>Ying Shen</strong></a>
</p>

  <h3 align="center"><a href="https://lif314.github.io/projects/realvlg_r1/">🌐Project page</a> 
  | <a href="https://arxiv.org/abs/2603.14880">📝Paper(arXiv)</a> | <a href="https://github.com/lif314/RealVLG-R1">💻Code </a>
  </h3>
  <div align="center"></div>
</p>

## Sample

Each data sample is annotated as follows:
```json
[
  {
    "image_name": "",
    "image_path": "",
    "object_id": "",
    "mask_path": "",
    "description": "",
    "label": "", # short description
    "bbox": [x1, y1, x2, y2],
    "grasps": [
      [x0,y0,x1,y1,x2,y2,x3,y3],
      ...
    ],
    "contact_points": [
    [x1,y1, x2, y2],
    ...
    ]
  }
]
```

The definition diagrams of bbox and grasp are shown in the figure below:
![](./assets/anno_demo.png)

## Usage

Download the dataset and extract `xxx_VLG.zip`. In each `xxx_VLG` folder, run `python metadata_viewer.py` to view the data formatting. The left/right keys switch between different objects in the same image, and the up/down keys switch between images. The visualization of different data subsets is shown below:

| Subdata | Cornell_VLG | VMRD_VLG | OCID_VLG | GraspNet_VLG | Jacquard_VLG |
|---------|-------------|----------|----------|--------------|---------------|
| Demo    | ![](./assets/cornell.png) | ![](./assets/vmrd.png) | ![](./assets/ocid.png) | ![](./assets/graspnet.png) | ![](./assets/jacquard.png) |

For more detailed data loading, please refer to `metadata_viewer.py`.

> Note: ``Jacquard_VLG`` is a simulated dataset not discussed in the paper. Its language annotations are derived from ShapeNetSem category labels.

## License

We thank all previous work. If you use this dataset, please cite the relevant work and comply with their licenses.

- [Cornell](https://www.kaggle.com/datasets/oneoneliu/cornell-grasp)
- [VMRD](https://opendatalab.com/OpenDataLab/VMRD)
- [OCID-Grasp](https://github.com/stefan-ainetter/grasp_det_seg_cnn)
- [GraspNet](https://graspnet.net/)
- [Jacquard](https://jacquard.liris.cnrs.fr/)
