---
title: P-MMEval
canonical_url: "https://www.modelscope.cn/datasets/Qwen/P-MMEval"
md_url: "https://www.modelscope.cn/datasets/Qwen/P-MMEval.md"
repository: Qwen/P-MMEval
chinese_name: P-MMEval
last_updated: 2024-12-10
license: "Apache License 2.0"
storage_size: "83 MB"
downloads: 4714
stars: 0
---

# P-MMEval

> P-MMEval - Qwen 在 ModelScope 开源的数据集。P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs

Qwen/P-MMEval 是 ModelScope 魔搭社区上的数据集，存储大小 83 MB，采用 Apache License 2.0 许可。

- **Repository**: Qwen/P-MMEval
- **License**: Apache License 2.0
- **Storage size**: 83 MB
- **Downloads**: 4714
- **Stars**: 0
- **Last updated**: 2024-12-10

Source: https://www.modelscope.cn/datasets/Qwen/P-MMEval

---

# P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs

## Introduction

We introduce a multilingual benchmark, P-MMEval, covering effective fundamental and capability-specialized datasets. We extend the existing benchmarks, ensuring consistent language coverage across all datasets and providing parallel samples among multiple languages, supporting up to 10 languages from 8 language families (i.e., en, zh, ar, es, ja, ko, th, fr, pt, vi). As a result, P-MMEval facilitates a holistic assessment of multilingual capabilities and comparative analysis of cross-lingual transferability.

## Supported Languages
- Arabic
- Spanish
- French
- Japanese
- Korean
- Portuguese
- Thai
- Vietnamese
- English
- Chinese

## Supported Tasks
<img src="https://cdn-uploads.huggingface.co/production/uploads/64abba3303cd5dee2efa6ee9/adic-93OnhRoSIk3P2VoS.png" width="1200" />

## Main Results

The multilingual capabilities of all models except for the LLaMA3.2 series improve with increasing model sizes, as LLaMA3.2-1B and LLaMA3.2-3B exhibit poor instruction-following capabilities, leading to a higher failure rate in answer extraction. In addition, Qwen2.5 demonstrates a strong multilingual performance on understanding and capability-specialized tasks, while Gemma2 excels in generation tasks. Closed-source models generally outperform open-source models.

<img src="https://cdn-uploads.huggingface.co/production/uploads/64abba3303cd5dee2efa6ee9/dGpAuDPT53TDHEW5wFZWk.png" width="1200" />

## Citation

We've published our paper at [this link](https://arxiv.org/pdf/2411.09116). If you find this dataset is helpful, please cite our paper as follows:
```
@misc{zhang2024pmmevalparallelmultilingualmultitask,
      title={P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs}, 
      author={Yidan Zhang and Yu Wan and Boyi Deng and Baosong Yang and Haoran Wei and Fei Huang and Bowen Yu and Junyang Lin and Fei Huang and Jingren Zhou},
      year={2024},
      eprint={2411.09116},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2411.09116}, 
}
```

# Usage
You can use OpenCompass if you want to evaluate your LLMs on P-MMEval . We advice you to use vllm to accelerate the evaluation (requiring vllm installation):

```
# CLI
opencompass --models hf_internlm2_5_1_8b_chat --datasets pmmeval_gen -a vllm

# Python scripts
opencompass ./configs/eval_PMMEval.py
```
