---
title: cvt-13
canonical_url: "https://www.modelscope.cn/models/microsoft/cvt-13"
md_url: "https://www.modelscope.cn/models/microsoft/cvt-13.md"
repository: microsoft/cvt-13
last_updated: 2025-07-21
license: apache-2.0
pipeline_tag: image-classification
tasks:
  - image-classification
model_type:
  - cvt
architectures:
  - CvtForImageClassification
parameters: 20.0M
tensor_type:
  - F32
  - I64
library_name:
  - transformer
  - safetensors
  - tensorflow
  - pytorch
frameworks:
  - pytorch
downloads: 181
stars: 0
tags:
  - vision
  - image-classification
---

# cvt-13

> cvt-13 - microsoft 在 ModelScope 开源的模型。Convolutional Vision Transformer (CvT)

microsoft/cvt-13 是 ModelScope 魔搭社区上的 20.0M 参数image-classification模型，采用 apache-2.0 许可。

- **Repository**: microsoft/cvt-13
- **License**: apache-2.0
- **Tasks**: image-classification
- **Parameters**: 20.0M
- **Tags**: vision, image-classification
- **Downloads**: 181
- **Stars**: 0
- **Last updated**: 2025-07-21

Source: https://www.modelscope.cn/models/microsoft/cvt-13

---

# Convolutional Vision Transformer (CvT)

CvT-13 model pre-trained on ImageNet-1k at resolution 224x224. It was introduced in the paper [CvT: Introducing Convolutions to Vision Transformers](https://arxiv.org/abs/2103.15808) by Wu et al. and first released in [this repository](https://github.com/microsoft/CvT). 

Disclaimer: The team releasing CvT did not write a model card for this model so this model card has been written by the Hugging Face team.

## Usage

Here is how to use this model to classify an image of the COCO 2017 dataset into one of the 1,000 ImageNet classes:

```python
from transformers import AutoFeatureExtractor, CvtForImageClassification
from PIL import Image
import requests

url = 'http://images.cocodataset.org/val2017/000000039769.jpg'
image = Image.open(requests.get(url, stream=True).raw)

feature_extractor = AutoFeatureExtractor.from_pretrained('microsoft/cvt-13')
model = CvtForImageClassification.from_pretrained('microsoft/cvt-13')

inputs = feature_extractor(images=image, return_tensors="pt")
outputs = model(**inputs)
logits = outputs.logits
# model predicts one of the 1000 ImageNet classes
predicted_class_idx = logits.argmax(-1).item()
print("Predicted class:", model.config.id2label[predicted_class_idx])
```
