---
title: llava-1.5-7b-hf-traffic
canonical_url: "https://www.modelscope.cn/models/zhaokaikai/llava-1.5-7b-hf-traffic"
md_url: "https://www.modelscope.cn/models/zhaokaikai/llava-1.5-7b-hf-traffic.md"
repository: zhaokaikai/llava-1.5-7b-hf-traffic
last_updated: 2025-10-25
model_type:
  - llava
architectures:
  - LlavaForConditionalGeneration
parameters: 7.1B
tensor_type:
  - BF16
library_name:
  - safetensors
  - pytorch
frameworks:
  - pytorch
inference_backends:
  - "deploy_task vlm/text/emb"
  - "lmdeploy 0.9.1"
  - "lmdeploy_turbomind 0.9.1"
  - "sglang 0.5.2"
  - "vllm 0.9.2"
downloads: 25
stars: 0
---

# llava-1.5-7b-hf-traffic

> llava-1.5-7b-hf-traffic - zhaokaikai 在 ModelScope 开源的模型。LLaVA-1.5-7B-HF-Traffic

zhaokaikai/llava-1.5-7b-hf-traffic 是 ModelScope 魔搭社区上的 7.1B 参数机器学习模型，可用 deploy_task vlm/text/emb、lmdeploy 0.9.1、lmdeploy_turbomind 0.9.1 部署。

- **Repository**: zhaokaikai/llava-1.5-7b-hf-traffic
- **Parameters**: 7.1B
- **Inference backends**: deploy_task vlm/text/emb, lmdeploy 0.9.1, lmdeploy_turbomind 0.9.1, sglang 0.5.2, vllm 0.9.2
- **Downloads**: 25
- **Stars**: 0
- **Last updated**: 2025-10-25

Source: https://www.modelscope.cn/models/zhaokaikai/llava-1.5-7b-hf-traffic

---

# LLaVA-1.5-7B-HF-Traffic

**LLaVA-1.5-7B-HF-Traffic** is a multimodal model fine-tuned on the **MITS (Multimodal Intelligent Traffic Surveillance)** dataset for intelligent traffic surveillance scenarios.

- **Tasks:** recognition, counting, localization, background awareness, reasoning
- **Data:** 170,400 images + ~5M instruction-following VQA pairs from MITS
- **Modality:** Image + Text → Text
- **Domain:** traffic scenes (congestion, accidents, construction, smoke/fireworks, unusual weather, spills, etc.)

## Quick Links
- 📚 Dataset: [`zhaokaikai/Multimodal_Intelligent_Traffic_Surveillance`](https://www.modelscope.cn/datasets/zhaokaikai/Multimodal_Intelligent_Traffic_Surveillance)
- 💻 Usage & examples: please refer to the GitHub repo  
  **https://github.com/LifeIsSoSolong/Multimodal-Intelligent-Traffic-Surveillance-Dataset-Models**

## Intended Use
- Urban traffic monitoring, incident analysis, visual question answering for transportation management
- Research on ITS-specific multimodal reasoning and instruction following

## Model Inputs/Outputs
- **Input:** an image (traffic scene) + a natural language instruction/question
- **Output:** a natural language response (e.g., description, count, event reasoning)

## Training Summary
- Objective: instruction tuning on MITS traffic QA
- Backbone family: LLaVA-1.5 (7B)
- Notes: align vision-language features to traffic-centric concepts and events

## Limitations & Notes
- The model may make mistakes on rare objects or extreme weather/night scenes not well represented in training.
- Not a safety-critical system; human verification is required for real-world decisions.

## License
- Follow the licenses of this model and the MITS dataset as stated on their ModelScope pages.

## Citation
If you use this model or dataset, please cite:
```bibtex
@article{zhao2025mits,
  title   = {MITS: A large-scale multimodal benchmark dataset for Intelligent Traffic Surveillance},
  author  = {Zhao, Kaikai and Liu, Zhaoxiang and Wang, Peng and Wang, Xin and Ma, Zhicheng and Xu, Yajun and Zhang, Wenjing and Nan, Yibing and Wang, Kai and Lian, Shiguo},
  journal = {Image and Vision Computing},
  pages   = {105736},
  year    = {2025},
  publisher = {Elsevier}
}
```

## Contact
UnicomAI
