---
title: MINICPM-V26_TPU
canonical_url: "https://www.modelscope.cn/models/radxa/MINICPM-V26_TPU"
md_url: "https://www.modelscope.cn/models/radxa/MINICPM-V26_TPU.md"
repository: radxa/MINICPM-V26_TPU
last_updated: 2024-12-25
pipeline_tag: visual-question-answering
tasks:
  - visual-question-answering
library_name:
  - other
frameworks:
  - other
downloads: 45
stars: 0
---

# MINICPM-V26_TPU

> MINICPM-V26_TPU - radxa 在 ModelScope 开源的模型。MiniCPM-V 2.6, a multimodal LLM designed for edge devices. MiniCPM-V 2.6 TPU is the adaptation of the OpenBMB open-source MiniCPM-V 2.6 multimodal language model to the SG2300X chip series using the Sophon SDK. This…

- **Repository**: radxa/MINICPM-V26_TPU
- **Tasks**: visual-question-answering
- **Downloads**: 45
- **Stars**: 0
- **Last updated**: 2024-12-25

Source: https://www.modelscope.cn/models/radxa/MINICPM-V26_TPU

---

# MiniCPM-V 2.6 TPU


[MiniCPM-V 2.6](https://github.com/OpenBMB/MiniCPM-V), a multimodal LLM designed for edge devices. MiniCPM-V 2.6 TPU is the adaptation of the OpenBMB open-source [MiniCPM-V 2.6](https://github.com/OpenBMB/MiniCPM-V) multimodal language model to the SG2300X chip series using the Sophon SDK. This adaptation enables hardware-accelerated inference via local TPUs, allowing users to ask questions about the content of input images.

### TPU Configuration

**Recommended TPU Memory Settings:**  
NPU -> 7615MB, VPU -> 2360MB, VPP -> 2360MB. [How to modify?](https://docs.radxa.com/en/sophon/airbox/local-ai-deploy/ai-tools/memory_allocate)

## Application Deployment

- Clone the Repository

  ```bash
  git clone https://github.com/zifeng-radxa/LLM-TPU.git
  ```

- Download Quantized Models and Precompiled C++ Libraries

  This example provides a pre-quantized `minicpmv26_bm1684x_int4_seq1024.bmodel` and precompiled dynamic libraries.  
   Refer to [MiniCPM-V 2.6 Model Conversion](#minicpm-v-26-model-conversion) for converting models of different lengths.

   Refer to [MiniCPM-V 2.6 CPython Compilation](#minicpm-v-26-cpython-compilation) for building CPython binding files.

- Download Precompiled Models Using [git LFS](https://git-lfs.com/)
  Models are available on [ModelScope](https://modelscope.cn/models/radxa/MINICPM-V26_TPU).

  ```bash
  cd LLM-TPU/models/MiniCPM-V-2_6/python_demo
  git clone https://www.modelscope.cn/radxa/MINICPM-V26_TPU.git
  mv MINICPM-V26_TPU/* .
  ```

- Set Up the Environment

  **It is recommended to create a virtual environment** to avoid conflicts with other applications. Refer to [this guide](../ai-tools/virtualenv_usage) for virtual environment usage.

  ```bash
  python3 -m virtualenv .venv
  source .venv/bin/activate
  ```

- Install Dependencies

  ```bash
  pip3 install --upgrade pip
  pip3 install torch torchvision pillow transformers
  ```

- Set Environment Variables
  Ensure the path for `libbmlib.so` linked to `chat.cpython-38-aarch64-linux-gnu.so` is correct. Use the `ldd` command to verify.  
   If the path is incorrect, update it as follows:

  ```bash
  export LD_LIBRARY_PATH=LLM-TPU/models/MiniCPM-V-2_6/support/lib_soc:$LD_LIBRARY_PATH
  ```

- Start MiniCPM-V 2.6

  Terminal Mode

  ```bash
  python3 pipeline.py -m ./minicpmv26_bm1684x_int4_seq1024.bmodel
  ```

  `-m`: Specify the model path.


## MiniCPM-V 2.6 Model Conversion

Follow these steps to convert MiniCPM-V 2.6 models to different lengths and quantization of `bmodel`.

- Prepare Docker Environment on an x86 Workstation

  Refer to [TPU-MLIR Installation](https://docs.radxa.com/en/sophon/airbox/model-compile/tpu_mlir_env#environment-setup) for setup instructions.  
   Clone the repository:

  ```bash
  git clone https://github.com/zifeng-radxa/LLM-TPU.git
  ```

- Download the MiniCPM-V 2.6 Model
  From [ModelScope](https://modelscope.cn/models/OpenBMB/MiniCPM-V-2_6/summary).

- Create a Virtual Environment in the Work Directory

  ```bash
  cd LLM-TPU/models/MiniCPM-V-2_6/compile
  python3 -m virtualenv .venv
  source .venv/bin/activate
  pip3 install --upgrade pip
  pip3 install torch torchvision --index-url https://download.pytorch.org/whl/cpu
  pip3 install transformers_stream_generator einops tiktoken accelerate transformers==4.40.0 onnx
  ```

- Align Model Environment

  Copy `LLM-TPU/models/MiniCPM-V 2.6/compile/files/MiniCPM-V-2_6/modeling_qwen2.py` into the Transformers library.
  Ensure that the Transformers library is located within the `.venv` environment.

  ```bash
  cp files/MiniCPM-V-2_6/modeling_qwen2.py .venv/lib/python3.10/site-packages/transformers/models/qwen2
  cp files/MiniCPM-V-2_6/resampler.py YOUR_MiniCPM-V-2_6_PATH
  cp files/MiniCPM-V-2_6/modeling_navit_siglip.py YOUR_MiniCPM-V-2_6_PATH
  ```

- Generate ONNX File

  ```bash
  python3 export_onnx.py --model_path YOUR_MiniCPM-V-2_6_PATH --seq_length 1024 --device cpu --image_file ../python_demo/test0.jpg
  ```

- Generate BModel File

  Exit the virtual environment:

  ```bash
  deactivate
  ```

  Compile the model:

  ```bash
  ./compile.sh --mode int4 --name minicpmv26 --seq_length 1024
  ```

  - `--mode`: Quantization mode (int4, int8, bf16).
  - `--seq_length`: Sequence length (e.g., 512, 1024, 2048).

  :::tip  
   Model compilation takes over 1 hour and requires at least 64GB memory and 200GB storage.  
   :::

## MiniCPM-V 2.6 CPython Compilation

Precompiled files are included in the model package. To compile manually:

```bash
cd python_demo
pip3 install pybind11
cmake -B build && cmake --build build
cp ./build/*.so .
```
