---
title: video_annotation_pipeline
canonical_url: "https://www.modelscope.cn/datasets/MBZUAI/video_annotation_pipeline"
md_url: "https://www.modelscope.cn/datasets/MBZUAI/video_annotation_pipeline.md"
repository: MBZUAI/video_annotation_pipeline
last_updated: 2025-03-17
license: "Apache License 2.0"
storage_size: "55 GB"
downloads: 1026
stars: 0
---

# video_annotation_pipeline

> video_annotation_pipeline - MBZUAI 在 ModelScope 开源的数据集。👁️ Semi-Automatic Video Annotation Pipeline

MBZUAI/video_annotation_pipeline 是 ModelScope 魔搭社区上的数据集，存储大小 55 GB，采用 Apache License 2.0 许可。

- **Repository**: MBZUAI/video_annotation_pipeline
- **License**: Apache License 2.0
- **Storage size**: 55 GB
- **Downloads**: 1026
- **Stars**: 0
- **Last updated**: 2025-03-17

Source: https://www.modelscope.cn/datasets/MBZUAI/video_annotation_pipeline

---

# 👁️ Semi-Automatic Video Annotation Pipeline

---
## 📝 Description
Video-ChatGPT introduces the VideoInstruct100K dataset, which employs a semi-automatic annotation pipeline to generate 75K instruction-tuning QA pairs. To address the limitations of this annotation process, we present VCG+112K dataset developed through an improved annotation pipeline. Our approach improves the accuracy and quality of instruction tuning pairs by improving keyframe extraction, leveraging SoTA large multimodal models (LMMs) for detailed descriptions, and refining the instruction generation strategy.


<p align="center">
  <img src="video_annotation_pipeline.png" alt="Contributions">
</p>


## 💻 Download
To get started, follow these steps:
   ```
   git lfs install
   git clone https://huggingface.co/MBZUAI/video_annotation_pipeline
   ```


## 📚 Additional Resources
- **Paper:** [ArXiv](https://arxiv.org/abs/2406.09418).
- **GitHub Repository:** For training and updates: [GitHub - GLaMM](https://github.com/mbzuai-oryx/VideoGPT-plus).
- **HuggingFace Collection:** For downloading the pretrained checkpoints, VCGBench-Diverse Benchmarks and Training data, visit [HuggingFace Collection - VideoGPT+](https://huggingface.co/collections/MBZUAI/videogpt-665c8643221dda4987a67d8d).

## 📜 Citations and Acknowledgments

```bibtex
  @article{Maaz2024VideoGPT+,
      title={VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding},
      author={Maaz, Muhammad and Rasheed, Hanoona and Khan, Salman and Khan, Fahad Shahbaz},
      journal={arxiv},
      year={2024},
      url={https://arxiv.org/abs/2406.09418}
  }
