---
title: pixmo-cap
canonical_url: "https://www.modelscope.cn/datasets/allenai/pixmo-cap"
md_url: "https://www.modelscope.cn/datasets/allenai/pixmo-cap.md"
repository: allenai/pixmo-cap
last_updated: 2025-05-28
license: "Apache License 2.0"
storage_size: "1.0 GB"
downloads: 651
stars: 0
---

# pixmo-cap

> pixmo-cap - allenai 在 ModelScope 开源的数据集。PixMo-Cap PixMo-Cap is a dataset of very long (roughly 200 words on average), detailed captions. It can be used to pre-train and fine-tune vision-language models. PixMo-Cap was created by recording annotators speaking…

allenai/pixmo-cap 是 ModelScope 魔搭社区上的数据集，存储大小 1.0 GB，采用 Apache License 2.0 许可。

- **Repository**: allenai/pixmo-cap
- **License**: Apache License 2.0
- **Storage size**: 1.0 GB
- **Downloads**: 651
- **Stars**: 0
- **Last updated**: 2025-05-28

Source: https://www.modelscope.cn/datasets/allenai/pixmo-cap

---

# PixMo-Cap
PixMo-Cap is a dataset of very long (roughly 200 words on average), detailed captions.
It can be used to pre-train and fine-tune vision-language models. 
PixMo-Cap was created by recording annotators speaking about an image for 60-90 seconds and then using the [Claude large language model](https://claude.ai/) to turn the audio transcripts(s) into a long caption. 
The audio transcripts are also included.

PixMo-Cap is part of the [PixMo dataset collection](https://huggingface.co/collections/allenai/pixmo-674746ea613028006285687b) and was used to train the [Molmo family of models](https://huggingface.co/collections/allenai/molmo-66f379e6fe3b8ef090a8ca19)

Quick links:
- 📃 [Paper](https://molmo.allenai.org/paper.pdf)
- 🎥 [Blog with Videos](https://molmo.allenai.org/blog)


## Loading
```python
data = datasets.load_dataset("allenai/pixmo-cap", split="train")
```

## Data Format
Images are stored as URLs that will need to be downloaded separately. 

The `transcripts` fields contains one or more audio transcripts

The `caption` field contains the caption from the LLM.

## License
This dataset is licensed by ODC-BY-1.0. It is intended for research and educational use in accordance with Ai2's [Responsible Use Guidelines](https://allenai.org/responsible-use).
This dataset includes output data generated from Claude which are subject to Anthropic [terms of service](https://www.anthropic.com/legal/commercial-terms) and [usage policy](https://www.anthropic.com/legal/aup).
