---
title: TCM-Instruction-Tuning-ShizhenGPT
canonical_url: "https://www.modelscope.cn/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT"
md_url: "https://www.modelscope.cn/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.md"
repository: FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT
last_updated: 2025-08-25
license: apache-2.0
storage_size: "91 GB"
downloads: 781
stars: 0
---

# TCM-Instruction-Tuning-ShizhenGPT

> TCM-Instruction-Tuning-ShizhenGPT - FreedomIntelligence 在 ModelScope 开源的数据集。This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data,…

FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT 是 ModelScope 魔搭社区上的数据集，存储大小 91 GB，采用 apache-2.0 许可。

- **Repository**: FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT
- **License**: apache-2.0
- **Storage size**: 91 GB
- **Downloads**: 781
- **Stars**: 0
- **Last updated**: 2025-08-25

Source: https://www.modelscope.cn/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT

---

# <span>📚 Introduction</span>

This dataset is a fine-tuning dataset for [ShizhenGPT](https://github.com/FreedomIntelligence/ShizhenGPT), a multimodal LLM for **Traditional Chinese Medicine (TCM)**. We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.

For details, see our [paper](https://arxiv.org/abs/2508.14706) and [GitHub repository](https://github.com/FreedomIntelligence/ShizhenGPT).

# <span>📊 Dataset Overview</span>

The open-sourced fine-tuning dataset consists of three parts:

|                                      | Modality                       | Data Quantity |
| ------------------------------------ | ------------------------------ | ------------- |
| TCM Text Instructions   | 📝 Text                        | 87K           |
| TCM Visual Instructions | 📝 Text, 👁️ Visual            | 67K           |
| TCM Speech Instructions  | 📝 Text, 👁️ Visual, 🎙️ Audio | 91K           |

> ⚠️ Note: Since TCM signal datasets, such as pulse and smell, involve private information, we recommend users download them from the corresponding paper.

# <span>📖 Citation</span>
If you find our data useful, please consider citing our work!
```
@misc{chen2025shizhengptmultimodalllmstraditional,
      title={ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine}, 
      author={Junying Chen and Zhenyang Cai and Zhiheng Liu and Yunjin Yang and Rongsheng Wang and Qingying Xiao and Xiangyi Feng and Zhan Su and Jing Guo and Xiang Wan and Guangjun Yu and Haizhou Li and Benyou Wang},
      year={2025},
      eprint={2508.14706},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2508.14706},
}
```
