---
title: "ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks"
canonical_url: "https://www.modelscope.cn/papers/124404"
md_url: "https://www.modelscope.cn/papers/124404.md"
arxiv_id: 2503.06885
published: 2025-03-10
last_updated: 2025-03-10
authors:
  - "Yan Yang"
  - "Dongxu Li"
  - "Haoning Wu"
  - "Bei Chen"
  - "Liu Liu"
  - "Liyuan Pan"
  - "Junnan Li"
model_name: ProBench
model_developer: "华为, 北京理工大学, Salesforce AI Research"
domain:
  - "自然语言处理"
  - "计算机视觉"
  - "深度学习"
  - "机器学习"
type:
  - "自然语言处理"
  - "计算机视觉"
  - "深度学习"
  - "机器学习"
  - "Computer Vision and Pattern Recognition (cs.CV)"
arxiv_url: "https://arxiv.org/abs/2503.06885"
pdf_url: "https://arxiv.org/pdf/2503.06885.pdf"
---

# ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks

> Solving expert-level multimodal tasks is a key milestone towards general intelligence. As the capabilities of multimodal large language models (MLLMs) continue to improve, evaluation of such advanced multimodal intelligence becomes necessary yet challenging.…

「ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks」是 ModelScope 魔搭社区收录的论文，arXiv 2503.06885，作者为 Yan Yang, Dongxu Li, Haoning Wu et al.，发表于 2025-03-10，属于 自然语言处理、计算机视觉、深度学习 领域。

- **ArXiv**: 2503.06885
- **Published**: 2025-03-10
- **Authors**: Yan Yang, Dongxu Li, Haoning Wu, Bei Chen, Liu Liu, Liyuan Pan, Junnan Li
- **Model**: ProBench
- **Developer**: 华为, 北京理工大学, Salesforce AI Research
- **Domain**: 自然语言处理, 计算机视觉, 深度学习, 机器学习
- **ArXiv URL**: https://arxiv.org/abs/2503.06885
- **PDF**: https://arxiv.org/pdf/2503.06885.pdf

Source: https://www.modelscope.cn/papers/124404

---

> ProBench：多模态大语言模型的专业级开放任务基准测试

## 摘要

本文介绍了ProBench，一个用于评估多模态大语言模型（MLLMs）在专业工作场景中表现的基准测试。ProBench包含4000个高质量样本，涵盖10个任务领域和56个子领域，支持多达13轮对话，并涉及17种语言。该基准测试旨在评估MLLMs在开放性任务中的表现，特别是那些需要专业知识和高级推理能力的任务。与现有基准不同，ProBench更注重用户指令的遵循和人类偏好的对齐，这些是实际应用中的关键方面。通过MLLM-as-a-Judge框架，ProBench能够自动评估和排名24个最新的MLLMs，揭示了当前模型在视觉感知、文本理解、领域知识和高级推理方面的局限性。实验结果表明，尽管开源模型的表现接近专有模型，但ProBench仍提出了显著挑战，为未来多模态AI研究提供了宝贵方向。

## Abstract

Solving expert-level multimodal tasks is a key milestone towards general intelligence. As the capabilities of multimodal large language models (MLLMs) continue to improve, evaluation of such advanced multimodal intelligence becomes necessary yet challenging. In this work, we introduce ProBench, a benchmark of open-ended user queries that require professional expertise and advanced reasoning. ProBench consists of 4,000 high-quality samples independently submitted by professionals based on their daily productivity demands. It spans across 10 fields and 56 sub-fields, including science, arts, humanities, coding, mathematics, and creative writing. Experimentally, we evaluate and compare 24 latest models using MLLM-as-a-Judge. Our results reveal that although the best open-source models rival the proprietary ones, ProBench presents significant challenges in visual perception, textual understanding, domain knowledge and advanced reasoning, thus providing valuable directions for future multimodal AI research efforts.
