---
title: MiniGPT-v2
canonical_url: "https://www.modelscope.cn/models/damo/MiniGPT-v2"
md_url: "https://www.modelscope.cn/models/damo/MiniGPT-v2.md"
repository: damo/MiniGPT-v2
last_updated: 2023-12-15
license: "BSD 3-Clause License"
parameters: 6.7B
tensor_type:
  - F16
library_name:
  - safetensors
  - pytorch
frameworks:
  - pytorch
language:
  - en
domain:
  - multi-modal
downloads: 1263
stars: 3
tags:
  - transformer
  - "arxiv:2310.09478"
---

# MiniGPT-v2

> MiniGPT-v2 - damo 在 ModelScope 开源的模型。MiniGPT4-V2是什么 MiniGPT4-V2是视觉-语言多任务学习中以大语言模型作为统一接口的AI算法。可用于图文对话。 使用方法 modelscope环境准备 直接使用docker镜像。具体可参考环境安装

damo/MiniGPT-v2 是 ModelScope 魔搭社区上的 6.7B 参数机器学习模型，采用 BSD 3-Clause License 许可。

- **Repository**: damo/MiniGPT-v2
- **License**: BSD 3-Clause License
- **Parameters**: 6.7B
- **Tags**: transformer, arxiv:2310.09478
- **Downloads**: 1263
- **Stars**: 3
- **Last updated**: 2023-12-15

Source: https://www.modelscope.cn/models/damo/MiniGPT-v2

---

# MiniGPT4-V2是什么
MiniGPT4-V2是视觉-语言多任务学习中以大语言模型作为统一接口的AI算法。可用于图文对话。  
![模型结构](image.png)
# 使用方法
## modelscope环境准备
直接使用docker镜像。具体可参考[环境安装](https://modelscope.cn/docs/环境安装)
```bash
docker pull registry.cn-hangzhou.aliyuncs.com/modelscope-repo/modelscope:ubuntu20.04-cuda11.8.0-py38-torch2.0.1-tf2.13.0-1.9.5
```
## 启动镜像并进入容器
```bash
docker run -it --gpus all --network host registry.cn-hangzhou.aliyuncs.com/modelscope-repo/modelscope:ubuntu20.04-cuda11.8.0-py38-torch2.0.1-tf2.13.0-1.9.5 bash
```
## 模型准备
可以使用git下载模型到任意位置。
```bash
git clone https://www.modelscope.cn/damo/MiniGPT-v2.git
```
## 调用示例
下面代码中pipeline的model参数可以使用上一步clone下来的MiniGPT-v2路径，也可以写为“damo/MiniGPT-v2”。moda会把model参数的值当作一个路径，当这个路径不存在时会下模型文件到$MODELSCOPE_CACHE。
```python
from modelscope.pipelines import pipeline

# import os
# os.environ["low_resource"] = "False"

input = {
    "image": "examples_v2/float.png",
    # "image": "该字段开可以使用网络图片网址",
    "ask": "describe this image in detail.",
    "temperature": "0.6"
}
inference = pipeline('MiniGPT-v2', model="damo/MiniGPT-v2", model_revision="v1.0.4")
output = inference(input)
print(output)
```
# ms_wrapper.py 部分内容说明
## 模型加载
相关代码在MyCustomModel-->init_model方法里。这里可以看到模型配置文件、模型权重文件、gpu_id、等关键配置。更多配置请修改 $MINIGPTPATH（这个环境变量将在运行代码时自动设置） 中的源码。  
其中cfg.model_cfg.low_resource字段用于配置是否以8bit加载LLM模型，这里使用默认配置“True”，此时要求显卡支持int8推理。如果您的显卡不支持int8，请将该配置项设置为False，此时需要更多的显存。如果你没有显卡，需要约40G的内存。  
在我们的测试中，如果该配置项设置为True需要约12G（V100）显存，如果设置为False能跑满显存（V100和P100都能跑满）。
```python
args = Dict({
      "gpu_id": 0,
      "options": None,
      "cfg_path": os.path.join(os.getenv("MINIGPTPATH"), "eval_configs/minigptv2_eval.yaml")
})
args.update(Dict(kwargs))
cfg = Config(args)
cfg.gpu_id = args.gpu_id
device = 'cuda:{}'.format(args.gpu_id) if torch.cuda.is_available() else "cpu"
cfg.model_cfg.low_resource = True
# 根据环境变量修改“low_resource”的值。
if os.getenv("low_resource") in ["False", "false", "0"]:
    cfg.model_cfg.low_resource = False
elif os.getenv("low_resource") in ["True", "true", "1"]:
    cfg.model_cfg.low_resource = True
# 当只有cpu可用时low_resource为False
if device == "cpu":
    cfg.model_cfg.low_resource = False
cfg.model_cfg.llama_model = args.model.llama_model or os.path.join(self.model_dir, "weight/Llama-2-7b-chat-hf")
cfg.model_cfg.ckpt = args.model.ckpt or os.path.join(self.model_dir, "weight/checkpoint_stage3.pth")
```
## 模型输入
image: 输入模型的图片，可以是一个本地文件的路径，会在这个字符串前拼接 $MINIGPTPATH 。也可以是一个网络图片网址，会直接从网络加载一张图。  
ask：对话的文本内容。
```python
input = {
        "image": "examples_v2/float.png",
        "ask": "[grounding] describe this image in detail.",
        "temperature": 0.6
    }
```
# 引用
```
@article{chen2023minigptv2,
      title={MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning}, 
      author={Chen, Jun and Zhu, Deyao and Shen, Xiaoqian and Li, Xiang and Liu, Zechu and Zhang, Pengchuan and Krishnamoorthi, Raghuraman and Chandra, Vikas and Xiong, Yunyang and Elhoseiny, Mohamed},
      year={2023},
      journal={arXiv preprint arXiv:2310.09478},
}

@article{zhu2023minigpt,
  title={MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models},
  author={Zhu, Deyao and Chen, Jun and Shen, Xiaoqian and Li, Xiang and Elhoseiny, Mohamed},
  journal={arXiv preprint arXiv:2304.10592},
  year={2023}
}
```
