---
title: M3SciQA
canonical_url: "https://www.modelscope.cn/datasets/yale-nlp/M3SciQA"
md_url: "https://www.modelscope.cn/datasets/yale-nlp/M3SciQA.md"
repository: yale-nlp/M3SciQA
last_updated: 2025-01-29
license: "Apache License 2.0"
storage_size: "137 MB"
downloads: 1632
stars: 0
---

# M3SciQA

> M3SciQA - yale-nlp 在 ModelScope 开源的数据集。🧑‍🔬 M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark For Evaluating Foundatio Models

yale-nlp/M3SciQA 是 ModelScope 魔搭社区上的数据集，存储大小 137 MB，采用 Apache License 2.0 许可。

- **Repository**: yale-nlp/M3SciQA
- **License**: Apache License 2.0
- **Storage size**: 137 MB
- **Downloads**: 1632
- **Stars**: 0
- **Last updated**: 2025-01-29

Source: https://www.modelscope.cn/datasets/yale-nlp/M3SciQA

---

# 🧑‍🔬 M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark For Evaluating Foundatio Models 

**EMNLP 2024 Findings**

🖥️ [Code](https://github.com/yale-nlp/M3SciQA)


## Introduction 

![image/png](./figures/overview.png)

In the realm of foundation models for scientific research, current benchmarks predominantly focus on single-document, text-only tasks and fail to adequately represent the complex workflow of such research. These benchmarks lack the $\textit{multi-modal}$, $\textit{multi-document}$ nature of scientific research, where comprehension also arises from interpreting non-textual data, such as figures and tables, and gathering information across multiple documents. To address this issue, we introduce M3SciQA, a Multi-Modal, Multi-document Scientific Question Answering benchmark designed for a more comprehensive evaluation of foundation models. M3SciQA consists of 1,452 expert-annotated questions spanning 70 natural language processing (NLP) papers clusters, where each cluster represents a primary paper along with all its cited documents, mirroring the workflow of comprehending a single paper by requiring multi-modal and multi-document data. With M3SciQA, we conduct a comprehensive evaluation of 18 prominent foundation models. Our results indicate that current foundation models still significantly underperform compared to human experts in multi-modal information retrieval and in reasoning across multiple scientific documents. Additionally, we explore the implications of these findings for the development of future foundation models. 

## Main Results 

### Locality-Specific Evaluation 
![image/png](./figures/MRR.png)


### Detail-Specific Evaluation
![image/png](./figures/detail.png)

## Cite
