---
title: PPE-MBPP-Plus-Best-of-K
canonical_url: "https://www.modelscope.cn/datasets/lmarena-ai/PPE-MBPP-Plus-Best-of-K"
md_url: "https://www.modelscope.cn/datasets/lmarena-ai/PPE-MBPP-Plus-Best-of-K.md"
repository: lmarena-ai/PPE-MBPP-Plus-Best-of-K
last_updated: 2025-04-21
license: "Apache License 2.0"
storage_size: "9.6 MB"
downloads: 770
stars: 0
---

# PPE-MBPP-Plus-Best-of-K

> PPE-MBPP-Plus-Best-of-K - lmarena-ai 在 ModelScope 开源的数据集。This contains the MBPP-Plus correctness preference evaluation set for Preference Proxy Evaluations.

lmarena-ai/PPE-MBPP-Plus-Best-of-K 是 ModelScope 魔搭社区上的数据集，存储大小 9.6 MB，采用 Apache License 2.0 许可。

- **Repository**: lmarena-ai/PPE-MBPP-Plus-Best-of-K
- **License**: Apache License 2.0
- **Storage size**: 9.6 MB
- **Downloads**: 770
- **Stars**: 0
- **Last updated**: 2025-04-21

Source: https://www.modelscope.cn/datasets/lmarena-ai/PPE-MBPP-Plus-Best-of-K

---

# Overview

This contains the MBPP-Plus correctness preference evaluation set for Preference Proxy Evaluations.

The prompts are sampled from [MBPP-Plus](https://huggingface.co/datasets/evalplus/mbppplus).

This dataset is meant for benchmarking and evaluation, not for training.

[Paper](https://arxiv.org/abs/2410.14872)

[Code](https://github.com/lmarena/PPE)

# License

User prompts are licensed under Apache-2.0, and model outputs are governed by the terms of use set by the respective model providers.

# Citation

```
@misc{frick2024evaluaterewardmodelsrlhf,
      title={How to Evaluate Reward Models for RLHF}, 
      author={Evan Frick and Tianle Li and Connor Chen and Wei-Lin Chiang and Anastasios N. Angelopoulos and Jiantao Jiao and Banghua Zhu and Joseph E. Gonzalez and Ion Stoica},
      year={2024},
      eprint={2410.14872},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2410.14872}, 
}
```
